Anchor live broadcast data generation method and device, storage medium and electronic equipment
By combining face and voiceprint recognition technology, accurate identification of downcast anchors in live broadcasts with multiple people on the same screen has been solved, and the accuracy of live broadcast data generation has been improved.
Patent Information
- Application Number
- CN202510078468.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
In the live broadcast scenario of multiple people on the same screen, it is difficult to accurately identify the host on the broadcast, resulting in poor accuracy of live broadcast data generation.
By obtaining the live audio files corresponding to the current frame screen and the intercept frame interval in the real-time live broadcast screen, combining face and voiceprint recognition technology, the final anchor information of the anchor to be identified, and by comparing the anchor information of the front and back frame screens, the downcast anchor is accurately identified.
Accurate anchor identification and downcast detection in the live broadcast scenario of multiple people on the same screen, improving the accuracy of anchor live broadcast data generation.
Smart Images

Figure CN119996767A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, storage medium and electronic device for generating anchor live broadcast data. Background Art
[0002] With the popularity of live streaming and the continuous development of technology, the live streaming industry has gradually become an important part of the Internet culture and entertainment industry. Multi-person live streaming on the same screen has become a common form of live streaming, which is a form of live streaming in which multiple anchors can interact and display at the same time in the live studio. For live streaming in the live studio, the generation and analysis of the anchor's live streaming data is particularly important, which can not only help the platform, anchors and content creators better understand the audience's needs, optimize content strategies, improve user experience, and ultimately maximize commercial value. Among them, the anchor's live streaming data refers to a series of key indicators of the anchor during the live broadcast, which are used to measure and analyze the anchor's live streaming performance.
[0003] At present, the method usually adopted to generate the live broadcast data of the anchor is: based on the conventional dynamic tracking function in the live broadcast software, the anchors in the live broadcast screen are effectively distinguished and tracked, so as to determine the anchor who is currently broadcasting and the anchor who is offline in the live broadcast room, and further, the start and end time of the anchor who is offline is queried through the background, and finally the live broadcast data of the anchor who is offline during the live broadcast is determined. However, the dynamic tracking function in the existing live broadcast software is usually applicable to the scene of single-person live broadcast. Once multiple people live broadcast on the same screen in the live broadcast room, the recognition accuracy of the offline anchor is low, resulting in poor accuracy in generating the live broadcast data of the anchor. Summary of the invention
[0004] In order to improve the accuracy of anchor live broadcast data generation, the present application provides an anchor live broadcast data generation method, device, storage medium and electronic device.
[0005] In a first aspect of the present application, a method for generating live broadcast data of an anchor is provided, which specifically includes: Acquire a current frame of a real-time live broadcast screen in a target live broadcast room, wherein the current frame is a frame of the real-time live broadcast screen at a current time, and the current frame contains at least one anchor to be identified; Based on the live audio file corresponding to the current frame and the frame-cut interval, determine the first final anchor information of each anchor to be identified in the current frame, the frame-cut interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the last frame captured last time, and the duration between the current frame and the last frame is the periodic interval duration; Compare the first final host information corresponding to the current frame with the second final host information corresponding to each host in the previous frame to obtain a comparison result; According to the comparison result, the off-air host is determined and the live broadcast data of the off-air host is generated.
[0006] By adopting the above technical solution, after obtaining the current frame, the first and final anchor information of each anchor to be identified in the current frame is analyzed and determined based on the current frame and the live audio file, so as to comprehensively determine the anchor situation in the current live broadcast of the target live broadcast room from the dimensions of face and voiceprint recognition, and accurately track and identify the anchor in the live broadcast of multiple people on the same screen. Then, the first and final anchor information of each anchor is compared with the second and final anchor information of each anchor in the previous frame, so as to analyze and determine whether there is a reduction in anchor information, that is, whether there is a situation where the anchor has gone offline, so as to more accurately identify the anchor who has gone offline, thereby improving the accuracy of the anchor live broadcast data generation.
[0007] Optionally, determining the first and final anchor information of each anchor to be identified in the current frame based on the live audio file corresponding to the current frame and the frame-cut interval specifically includes: Based on the current frame, determine the first anchor information corresponding to each anchor to be identified; Determine the second host information of each host in the live audio file based on the live audio file corresponding to the frame clipping interval; Intersecting each of the first anchor information and each of the second anchor information, to obtain first final anchor information of each of the to-be-identified anchors in the current frame.
[0008] By adopting the above technical solution, the first anchor information of each anchor to be identified who is currently broadcasting live in the target live broadcast room is determined according to the current frame, and then the second anchor information of the anchor who has broadcast live in the target live broadcast room within the frame interval is determined according to the live broadcast audio file. Finally, the intersection of each first anchor information and each second anchor information is performed to obtain the final anchor information. Through the intersection method, the anchor information determined based on the two dimensions of voiceprint and face can be verified with each other, thereby more accurately determining the anchor information of each anchor to be identified in the current frame.
[0009] Optionally, determining the first anchor information corresponding to each anchor to be identified based on the current frame specifically includes: Based on the current frame, determining at least one face coordinate by a preset third-party face detection service; Based on the facial coordinates, the actual face of the corresponding host to be identified is captured from the current frame, and each actual face is compared with the host head portrait in a preset database to obtain a corresponding target face identification, wherein the database includes different host information and corresponding host head portraits and face sample identifications; The first anchor information corresponding to each target face identifier is matched from the database.
[0010] By adopting the above technical solution, face detection and recognition is performed on the current frame to determine the facial coordinates of each face contained in the current frame. In this way, the corresponding face can be accurately captured from the current frame according to the facial coordinates, which is convenient for face comparison with the anchor avatar in the database, so as to accurately determine the corresponding target face identification, and finally match the corresponding first anchor information according to the target face identification, thereby achieving more accurate determination of the anchor information of the anchor currently broadcasting live in the target live broadcast room from the perspective of face recognition.
[0011] Optionally, the determining, based on the live audio file corresponding to the frame-cut interval, the second host information of each host in the live audio file specifically includes: Based on the live audio file corresponding to the frame-cut interval, at least one target voiceprint identifier is determined through a preset third-party intelligent voice recognition service; The target anchor information corresponding to each target voiceprint identifier is matched from a preset database, and the target anchor information is determined as the second anchor information of each anchor in the live audio file, wherein the database includes different anchor information and corresponding voiceprint identifiers.
[0012] By adopting the above technical solution, through performing voiceprint recognition on the live audio file in the screenshot interval, the voiceprint identifier of the anchor who is broadcasting live in the target live room within the screenshot interval, that is, the target voiceprint identifier, can be matched. Finally, the corresponding target anchor information can be matched from the database according to the target voiceprint identifier, so as to more accurately determine the anchor information of the anchor who is broadcasting live in the target live room from the perspective of voiceprint recognition.
[0013] Optionally, determining the offline anchor and generating live broadcast data of the offline anchor according to the comparison result specifically includes: When the comparison result is that the anchor information is newly added or the anchor information remains unchanged, the step of obtaining the current frame picture in the real-time live broadcast picture of the target live broadcast room is repeatedly performed after the interval time, until the comparison result is that the anchor information is reduced, and the live broadcast data of the anchor who is broadcasting is determined to be offline; When the comparison result is that the anchor information is reduced, the second final anchor information other than the first final anchor information among all the second final anchor information is determined as the downcast anchor information, and the anchor corresponding to the downcast anchor information is determined as the downcast anchor; Determine the time node corresponding to the current frame as the target download time node, and select the target upload time node of the downloading anchor from the preset upload time record, wherein the upload time record includes anchor information of different anchors and the corresponding upload time nodes; The live broadcast data of the offline host is generated according to the first live broadcast data of the target live broadcast room at the target broadcast time node and the second live broadcast data of the target offline broadcast time node.
[0014] By adopting the above technical solution, if the comparison result is that the anchor information is newly added or the anchor information remains unchanged, it means that the anchor has not gone offline in the target live broadcast room between the time of the previous frame and the time of the current frame, and there is no need to generate the anchor's live broadcast data temporarily. Then, after the interval, the current frame is captured again to try to determine the offline anchor again; if the comparison result is that the anchor information is reduced, it means that there is an anchor currently offline in the target anchor room, then the time node corresponding to the current frame is determined as the time node of the offline anchor's offline broadcast, which can not only quickly lock the offline anchor, but also accurately determine the live broadcast data of the offline anchor during the live broadcast.
[0015] Optionally, the method further includes: When the comparison result is a new addition to the broadcast information, the time node corresponding to the current frame is determined as the broadcast time node; A mapping relationship is established between the newly added anchor information and the broadcast time node and added to the preset broadcast time record.
[0016] By adopting the above technical solution, when the comparison result is that the anchor information is newly added, it means that the target live broadcast room has no anchor going offline, but there is a new anchor going online. Then the time node corresponding to the current frame is determined as the online time node, and the anchor corresponding to the newly added anchor information is the newly online anchor. Furthermore, a mapping relationship is established between the newly added anchor information and the online time node, and added to the preset online time record, so as to facilitate the subsequent rapid determination of the online time of the anchor who is going offline, and then more accurately and comprehensively determine the live broadcast performance of the anchor who is going offline when broadcasting in the target live broadcast room.
[0017] Optionally, the method further includes: Obtaining the host information, host avatar and voice clip of at least one host to be recorded sent by the user's terminal; Based on the single anchor avatar, determine the face sample identifier of the corresponding anchor avatar through a preset third-party face detection service; Based on the single voice segment, determine the voiceprint identifier of the corresponding voice segment through a preset third-party intelligent voice recognition service; The anchor information, anchor avatar, face sample identification and voiceprint identification corresponding to each anchor to be recorded are stored in a preset database.
[0018] By adopting the above technical solution, the face sample identifier and voiceprint identifier of each host to be recorded are determined through the host's head portrait and voice clips input by the terminal, so as to realize the collection of the face and voiceprint of the host to be recorded. Furthermore, the host information, host head portrait, face sample identifier and voiceprint identifier corresponding to each host to be recorded are archived, so as to facilitate the subsequent recognition of the host's voiceprint and face in the live broadcast, and lock the host information of the host in the live broadcast.
[0019] In a second aspect of the present application, a host live broadcast data generating device is provided, which specifically includes: An information acquisition module is used to acquire a current frame of a real-time live broadcast screen in a target live broadcast room, wherein the current frame is a frame of the real-time live broadcast screen at a current time, and the current frame contains at least one anchor to be identified; The anchor identification module is used to determine the first and final anchor information of each anchor to be identified in the current frame based on the live audio file corresponding to the current frame and the frame-cut interval, wherein the frame-cut interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the last frame captured last time, and the duration between the current frame and the last frame is the periodic interval duration; An information comparison module, used to compare the first final anchor information corresponding to the current frame with the second final anchor information corresponding to each anchor in the previous frame to obtain a comparison result; The data generation module is used to determine the offline anchor and generate the live broadcast data of the offline anchor according to the comparison result.
[0020] By adopting the above technical solution, the information acquisition module obtains the current frame picture, and the anchor identification module determines the first and final anchor information of each anchor to be identified in the current frame picture based on the live audio file corresponding to the current frame picture and the frame cut-off interval. Then the information comparison module compares each first and final anchor information with each second and final anchor information to obtain a comparison result. Finally, the data generation module determines the off-air anchor and generates the live broadcast data of the off-air anchor according to the comparison result.
[0021] In a third aspect of the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is loaded and executed by a processor, the method steps described in any one of the first aspects are performed.
[0022] In a fourth aspect of the present application, an electronic device is provided, specifically comprising: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the processor is used to load and execute the computer program stored in the memory so that the electronic device performs the method as described in any one of the first aspects.
[0023] In summary, the present application includes at least one of the following beneficial technical effects: After obtaining the current frame, the first and final anchor information of each anchor to be identified in the current frame is analyzed and determined based on the current frame and the live audio file, so as to comprehensively determine the anchor situation in the current live broadcast of the target live broadcast room from the dimensions of face and voiceprint recognition, and accurately track and identify the anchor in the live broadcast of multiple people on the same screen. Then, the first and final anchor information of each anchor is compared with the second and final anchor information of each anchor in the previous frame, so as to analyze and determine whether there is a reduction in anchor information, that is, whether there is a situation where the anchor has gone offline, so as to more accurately identify the anchor who has gone offline, thereby improving the accuracy of the anchor live broadcast data generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of a method for generating live broadcast data of an anchor provided in an embodiment of the present application; Figure 2 It is a flowchart of another method for generating live broadcast data of an anchor provided in an embodiment of the present application; Figure 3 It is a structural diagram of a device for generating live broadcast data provided by an embodiment of the present application; Figure 4 It is a structural diagram of another anchor live broadcast data generating device provided in an embodiment of the present application.
[0025] Explanation of the accompanying drawings: 11. Information acquisition module; 12. Anchor identification module; 13. Information comparison module; 14. Data generation module; 15. Time recording module; 16. Data entry module. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0027] In the description of the embodiments of the present application, words such as "illustrative", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "illustrative", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "illustrative", "for example" or "for example" is intended to present related concepts in a concrete way.
[0028] In the description of the embodiments of the present application, the term "and / or" is only a kind of association relationship describing the associated objects, indicating that there may be three kinds of relationships, for example, A and / or B, which can represent: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0029] See also Figure 1 The embodiment of the present application discloses a flowchart of a method for generating live broadcast data of an anchor, which can be implemented by a computer program or run on an anchor live broadcast data generating device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application, specifically including: S101: Acquire the current frame of the real-time live broadcast picture of the target live broadcast room.
[0030] Specifically, in an embodiment of the present application, the target live broadcast room is a live broadcast room for live broadcasting. In other embodiments, the target live broadcast room can also be a live broadcast room mainly for leisure and entertainment. The current frame is the frame of the real-time live broadcast picture at the current time. The current frame contains at least one anchor to be identified, that is, the anchor who is currently broadcasting live in the target live broadcast room. The real-time live broadcast picture is the real-time picture of the anchor in the target live broadcast room. The picture frame is a single image picture of the smallest unit in a video or animation, specifically a single frame image in a real-time live broadcast video.
[0031] Furthermore, the execution subject of a method for generating live broadcast data of an anchor disclosed in an embodiment of the present application can be a server, and the server is wirelessly connected to the user's terminal. The terminal can be a personal computer or a tablet computer, and a client for generating live broadcast data of the anchor is installed in the terminal. The server can be a background server of the client, and can be an independent physical server or a server cluster composed of multiple physical servers. In addition, the implementation scenario of the present application is: if the relevant person in charge of the enterprise needs to determine the live broadcast data of the anchor from the start of the broadcast to the end of the broadcast, wherein the live broadcast data refers to a series of key indicators of the anchor during the live broadcast, which are used to measure and analyze the live broadcast performance of the anchor, and the live broadcast data includes but is not limited to the cumulative number of viewers, the number of transactions, the number of likes and the number of comments during the live broadcast, then click to open the above client in the terminal, click the "Generate Live Data in Real Time" button, and capture the anchor who is broadcasting offline in the target live broadcast room in real time, and generate the live broadcast data of the anchor according to the live broadcast data of the anchor who is broadcasting offline from the start of the broadcast to the end of the broadcast.
[0032] Furthermore, a feasible way to obtain the current frame is: at intervals of a specific duration, the current frame is captured from the video of the real-time live broadcast screen of the target live broadcast room through the OpenCV tool. In other embodiments, the current frame can also be captured through the imageio tool. Among them, the specific duration is a periodic interval duration, and the interval duration is 10 seconds. In other embodiments, the interval duration can be 5 seconds. Exemplarily, the time node of the last capture of the real-time live broadcast screen frame is 10:20:10, then when the current time is 10:20:20, the current frame of the real-time live broadcast screen is captured, that is, the current frame.
[0033] In other embodiments, before officially starting to generate the live broadcast data of the anchor in real time, the anchor information, anchor avatar and voice clip of at least one anchor to be recorded sent by the user's terminal are obtained, wherein the anchor avatar is an image containing the anchor's face, and the voice clip is a voice of a preset duration recorded by the anchor in advance, and the preset duration can be 30 seconds or 20 seconds. The anchors to be recorded are all anchors participating in the live broadcast work of the target live broadcast room. Further, the preset third-party face detection service is called, and the corresponding face sample is added to the anchor avatar of each anchor to be recorded for subsequent face comparison, and the corresponding face sample identifier is generated. Then the preset third-party intelligent voice recognition service and sound registration interface are called to determine the voiceprint identifier corresponding to each anchor to be recorded. Finally, the anchor information, anchor avatar, face sample identifier and voiceprint identifier corresponding to each anchor to be recorded are mapped and stored in a preset database. Among them, the third-party face detection service refers to a service provided by an independent third-party organization for detecting and verifying facial images. The third-party intelligent voice recognition service refers to a service provided by an independent third-party organization for voiceprint recognition.
[0034] S102: Determine first and final anchor information of each anchor to be identified in the current frame based on the live audio file corresponding to the current frame and the frame-cutting interval.
[0035] Specifically, the frame interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the previous frame captured last time, and the duration between the current frame and the previous frame is the periodic interval duration. Through Web Real-Time Communication (WebRTC) technology, the video data and audio data of the target live broadcast room can be obtained in real time, and the audio data between the time node corresponding to the previous frame and the time node corresponding to the current frame can be filtered out from the acquired audio data, that is, the live audio file. Since the current frame contains the image of the anchor and the live audio file contains the sound of the anchor during the live broadcast, in the embodiment of the present application, by combining the current frame and the live audio file, the situation of the anchor who is currently broadcasting in the target live broadcast room can be analyzed and determined more accurately and reasonably, that is, the first and final anchor information of each anchor to be identified in the current frame is determined. Among them, the anchor information is the anchor ID, which is equivalent to the anchor's "identity card number" and is used to identify and distinguish different anchors. Furthermore, a feasible way to determine the first and final anchor information is: based on the current frame, determine the first anchor information of each anchor to be identified, and the specific process is: submit the current frame to a preset third-party face detection service to detect the face in the current frame, and determine the precise location of each face and the key point information, and finally obtain the face coordinates of each anchor face to be identified in the current frame. Among them, face coordinates refer to a set of pixel coordinates used to locate and identify specific points on the face. In face recognition and analysis technology, face coordinates play a vital role and can help the system accurately find the location of the face.
[0036] Furthermore, through a third-party face detection service, according to each face coordinate, the actual face of the corresponding host to be identified is captured from the current frame, and then the actual face is compared with the host avatar in the database to obtain a comparison result, which includes the error recognition rate and confidence when the actual face is compared with the host avatar. The error recognition rate is compared with a preset error recognition rate threshold, and the confidence is compared with a preset confidence threshold. The error recognition rate threshold is set to 1 / 1000, and the confidence threshold is set to 61%. In other embodiments, it can also be set to other reasonable values. If the error recognition rate is lower than the error recognition rate threshold, and the confidence is higher than the confidence threshold, it is determined that the corresponding actual face and the host avatar are the same person, and then the face sample identifier to which the corresponding host avatar belongs is determined as the target face identifier of the actual face. Finally, the host information corresponding to the target face identifier is matched from the database, that is, the first host information. Among them, the False Acceptance Rate (FAR) refers to the proportion of non-target faces mistakenly identified as target faces by the system. The lower the false recognition rate, the more accurate the matching result. The confidence level refers to the system's trust in the recognition result, usually expressed as a percentage or score. The higher the confidence level, the more accurate the matching result.
[0037] Furthermore, based on the live audio file corresponding to the frame cut interval, the second anchor information of the anchor who made the sound in the live audio file is determined. The specific determination process is: calling the preset third-party intelligent speech recognition service to perform voiceprint recognition on this live audio file. In the process, the approximate ambient sound waveform audio track and the irregular noise waveform audio track are eliminated through the intelligent track segmentation algorithm. At the same time, the timestamp of the live audio file is verified to improve the voiceprint matching degree. This is a prior art and will not be repeated here. Then determine the target voiceprint identification corresponding to each sound contained in this live audio file. Finally, match the target anchor information corresponding to each target voiceprint identification from the database, and then determine the second anchor information of each anchor involved in the live audio file.
[0038] Furthermore, all the first anchor information and all the second anchor information are intersected to obtain the final anchor information of each anchor to be identified in the current frame, that is, the first final anchor information. By means of intersection, the anchor information determined based on the two dimensions of voiceprint and face can be verified with each other, thereby more accurately determining the anchor information of each anchor to be identified in the current frame.
[0039] S103: Compare each first final anchor information corresponding to the current frame with the second final anchor information corresponding to each anchor in the previous frame to obtain a comparison result.
[0040] Specifically, after the first final anchor information corresponding to the current frame is determined, the second final anchor information corresponding to each anchor in the previous frame of the current frame is determined. The determination method can refer to step S103 and will not be repeated here. Then, each first final anchor information is compared with each second final anchor information to obtain a comparison result. In the embodiment of the present application, the comparison result is that the anchor information is added, the anchor information is unchanged, or the anchor information is reduced. Specifically, if a certain anchor information exists in each second final anchor information, but does not exist in each first final anchor information, the comparison result is that the anchor information is reduced, indicating that there is currently an anchor going offline; if a certain anchor information does not exist in each second final anchor information, but exists in each first final anchor information, the comparison result is that the anchor information is added, indicating that there is currently an anchor going online; if each first final anchor information is consistent with each second final anchor information, it means that the anchor of the current live broadcast has not changed.
[0041] S104: According to the comparison result, determine the off-air host and generate live broadcast data of the off-air host.
[0042] Specifically, in the embodiment of the present application, if the comparison result is that the anchor information is reduced, then the anchor information that is reduced in each first anchor information compared to each second anchor information is determined as the anchor information of the offline anchor. Then, the live broadcast data of the target live broadcast room at the time node corresponding to the current frame is obtained through a preset third-party live broadcast analysis tool, wherein the third-party live broadcast analysis tool can be a Socialbakers tool. Then, the time node of the offline anchor's broadcast is filtered out from the preset broadcast time record, wherein the broadcast time record includes the anchor information of different anchors and the corresponding broadcast time node. At the same time, the live broadcast data of the time node is obtained. Based on the live broadcast data of the target live broadcast room at the above two time nodes, the live broadcast data of the downcast anchor in the target live broadcast room is determined. A feasible determination method is: the cumulative number of viewers in the target live broadcast room when the downcast anchor is off-broadcast minus the cumulative number of viewers in the target live broadcast room when the downcast anchor is on-broadcast, to obtain the newly added number of viewers; the number of transactions in the target live broadcast room when the downcast anchor is off-broadcast minus the number of transactions in the target live broadcast room when the downcast anchor is on-broadcast, to obtain the newly added number of transactions; the number of likes in the target live broadcast room when the downcast anchor is off-broadcast minus the number of likes in the target live broadcast room when the downcast anchor is on-broadcast, to obtain the newly added number of likes. Finally, the above-mentioned newly added number of viewers, newly added number of transactions, and newly added number of likes are determined as the live broadcast data of the downcast anchor during the live broadcast, so as to intuitively reflect its live broadcast performance during the live broadcast. In other embodiments, a video clip of the downcast anchor from the time of broadcasting to the time of downcasting can be intercepted from the live video of the target live broadcast room, and the video clip is associated with the live broadcast data of the downcast anchor.
[0043] See also Figure 2 The present application embodiment discloses a flowchart of another method for generating live broadcast data of an anchor, which can be implemented by a computer program or run on an anchor live broadcast data generating device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application, specifically including: S201: Acquire the current frame of the real-time live broadcast screen of the target live broadcast room.
[0044] S202: Determine first and final host information of each host to be identified in the current frame based on the live audio file corresponding to the current frame and the frame-cut interval.
[0045] S203: Compare each first final anchor information corresponding to the current frame with the second final anchor information corresponding to each anchor in the previous frame to obtain a comparison result.
[0046] For details, please refer to steps S101-S103, which will not be described in detail here.
[0047] S204: When the comparison result is that the anchor information is newly added or the anchor information remains unchanged, repeatedly execute the step of obtaining the current frame picture in the real-time live broadcast picture of the target live broadcast room until the comparison result is that the anchor information is reduced, and determine the live broadcast data of the offline anchor.
[0048] Specifically, in the embodiment of the present application, if the comparison result is that the anchor information is newly added or the anchor information remains unchanged, it means that the anchor has not gone offline in the target live broadcast room between the time of the previous frame and the time of the current frame, and there is no need to generate the anchor's live broadcast data temporarily. Then, after the interval time, the step of obtaining the current frame in the real-time live broadcast screen of the target live broadcast room is repeated, and the current frame of the real-time live broadcast screen is re-captured. Then, according to the live audio file corresponding to the current frame and the frame capture interval, the first and final anchor information of each anchor to be identified in the current frame is determined, and compared with the second and final anchor information of each anchor in the previous frame. When the comparison result shows that the anchor information is reduced, the live broadcast data of the offline anchor is re-determined. For details, please refer to steps S101-S104, which will not be repeated here.
[0049] In other embodiments, when the comparison result is that the anchor information is newly added, it means that there is no anchor going offline in the target live broadcast room, but a new anchor is going online. In this case, the time node corresponding to the current frame is determined as the online broadcast time node, and the anchor corresponding to the newly added anchor information is the newly online anchor. Furthermore, a mapping relationship is established between the newly added anchor information and the online broadcast time node, and the mapping relationship is added to the preset online broadcast time record, so as to facilitate the subsequent rapid determination of the online broadcast time of the anchor who is going offline, and then more accurately and comprehensively determine the live broadcast performance of the anchor who is going offline when broadcasting in the target live broadcast room.
[0050] S205: When the comparison result shows that the anchor information is reduced, the second final anchor information except the first final anchor information among all the second final anchor information is determined as the downcast anchor information, and the anchor corresponding to the downcast anchor information is determined as the downcast anchor.
[0051] S206: Determine the time node corresponding to the current frame as the target downloading time node, and select the target broadcasting time node of the offline broadcaster from the preset broadcasting time record.
[0052] S207: Generate live broadcast data of the offline host according to the first live broadcast data of the target live broadcast room at the target online broadcast time node and the second live broadcast data at the target offline broadcast time node.
[0053] Specifically, when the comparison result shows that the anchor information is reduced, it means that there is an anchor currently broadcasting in the target anchor room, then the second final anchor information except the first final anchor information in all the second final anchor information is determined as the broadcast anchor information, and the anchor corresponding to this broadcast anchor information is determined as the broadcast anchor. Further, the time node corresponding to the current frame is determined as the target broadcast time node, and the target broadcast time node of the broadcast anchor is filtered from the preset broadcast time record. Then, based on the first live broadcast data of the target live broadcast room at the target broadcast time node and the second live broadcast data at the target broadcast time node, the live broadcast data of the broadcast anchor is generated. For details, please refer to steps S101-S104, which will not be repeated here.
[0054] In another embodiment, for a single anchor participating in the live broadcast of the target live broadcast room, the first transaction quantity of each historical product category during the historical live broadcast is counted, and the first number of historical product categories are selected from each historical product category in the order of the first transaction quantity from large to small to determine as the target product category, that is, the product category that is easy to trade. Then, the second transaction quantity of a single target product category in each historical live broadcast time period is obtained, and the second number of historical live broadcast time periods are selected from each historical live broadcast time period in the order of the second transaction quantity from large to small to determine as the target live broadcast time period corresponding to the target product category, that is, the live broadcast time period that is easy to trade. Further, the first weight of each target product category and the second weight of each corresponding target live broadcast time period, the first weight is the ratio of the first transaction quantity of each target product category to the sum of the first transaction quantities of all target product categories, and the second weight is the ratio of the second transaction quantity of a single target live broadcast time period corresponding to the target product category to the sum of the second transaction quantities of all corresponding target live broadcast time periods.
[0055] Furthermore, the host to be broadcasted in the next live broadcast period after the current one and at least one category to be sold in the next live broadcast period are obtained from the preset host live broadcast schedule. If the next live broadcast period exists in each target live broadcast period corresponding to the target category of the host to be broadcasted, then the target category to be sold is determined as the key category to be sold. The categories to be sold in each key category to be sold are determined as the final categories to be sold, and the first product of the first weight of each final category to be sold and the second weight of the corresponding next live broadcast period is calculated, and each first product is summed to obtain the sum of the first products. The larger the sum of the first products, the better the transaction performance of the host to be broadcasted in the subsequent live broadcast. Finally, the sum of the first products is compared with the preset product and threshold. If the sum of the first products is greater than the product and threshold, it means that the host to be broadcasted has a high probability of having a good transaction performance in the subsequent live broadcast, which means that it is more reasonable to arrange the host to be broadcasted to broadcast in the next live broadcast period, thereby realizing the rationality verification of the host live broadcast schedule. If the sum of the first products is not greater than the product sum threshold, it means that arranging the host to be broadcasted in the next live broadcast time period is not reasonable and the transaction effect is poor. The live broadcast time period of the host to be broadcasted needs to be rearranged, and the live broadcast time period after the next live broadcast time period is determined as the reference live broadcast time period. The planned product category within a single reference live broadcast time period is determined, and the second product of the first weight of the planned product category and the second weight of the corresponding reference live broadcast time period is calculated and summed to obtain the sum of the second products of the corresponding reference live broadcast time period. The larger the sum of the second products is, the better the live broadcast effect of the host to be broadcasted in the reference live broadcast time period. The maximum sum of the second products is selected from the sums of the second products, and the reference live broadcast time period corresponding to the maximum sum of the second products is adjusted to the live broadcast time period matching the host to be broadcasted.
[0056] The implementation principle of a method for generating live broadcast data of an anchor in an embodiment of the present application is as follows: after obtaining the current frame, based on the current frame and the live broadcast audio file, the first and final anchor information of each anchor to be identified in the current frame is analyzed and determined, thereby realizing the comprehensive determination of the anchor situation in the current live broadcast of the target live broadcast room from the dimensions of face and voiceprint recognition, and accurately tracking and identifying the anchor in the live broadcast of multiple people on the same screen. Then, each first and final anchor information is compared with the second and final anchor information of each anchor in the previous frame, so as to analyze and determine whether there is a situation of reduced anchor information, that is, whether there is a situation of anchor going offline, so as to more accurately identify the anchor who is going offline, thereby improving the accuracy of anchor live broadcast data generation.
[0057] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0058] See also Figure 3, is a schematic diagram of the structure of the anchor live broadcast data generation device provided in the embodiment of the present application. The anchor live broadcast data generation device can be implemented as all or part of the device through software, hardware or a combination of both. The device includes an information acquisition module 11, an anchor identification module 12, an information comparison module 13 and a data generation module 14.
[0059] The information acquisition module 11 is used to acquire the current frame of the real-time live broadcast screen of the target live broadcast room, where the current frame is the frame of the real-time live broadcast screen at the current time, and the current frame contains at least one anchor to be identified; The anchor identification module 12 is used to determine the first and final anchor information of each anchor to be identified in the current frame based on the live audio file corresponding to the current frame and the frame-cut interval, the frame-cut interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the last frame captured last time, and the duration between the current frame and the last frame is the periodic interval duration; The information comparison module 13 is used to compare the first final anchor information corresponding to the current frame with the second final anchor information corresponding to each anchor in the previous frame to obtain a comparison result; The data generating module 14 is used to determine the host who is going offline and generate the live broadcast data of the host who is going offline according to the comparison result.
[0060] Optionally, the anchor identification module 12 is specifically used for: Based on the current frame, determine the first anchor information corresponding to each anchor to be identified; Determine the second host information of each host in the live audio file based on the live audio file corresponding to the frame cut interval; The first anchor information is intersected with the second anchor information to obtain the first and final anchor information of each anchor to be identified in the current frame.
[0061] Optionally, the anchor identification module 12 is specifically used for: Based on the current frame, determine at least one face coordinate through a preset third-party face detection service; Based on the coordinates of each face, the actual face of the corresponding host to be identified is captured from the current frame, and each actual face is compared with the host avatar in the preset database to obtain the corresponding target face identification. The database includes different host information and corresponding host avatars and face sample identifications; Match the first anchor information corresponding to each target face identifier from the database.
[0062] Optionally, the anchor identification module 12 is specifically used for: Based on the live audio file corresponding to the frame-cut interval, at least one target voiceprint identifier is determined through a preset third-party intelligent voice recognition service; The target anchor information corresponding to each target voiceprint identifier is matched from a preset database, and the target anchor information is determined as the second anchor information of each anchor in the live audio file. The database includes different anchor information and corresponding voiceprint identifiers.
[0063] Optionally, the data generating module 14 is specifically used for: When the comparison result is that the anchor information is newly added or the anchor information remains unchanged, the step of obtaining the current frame of the real-time live broadcast screen of the target live broadcast room is repeatedly performed after the interval time, until the comparison result is that the anchor information is reduced, and the live broadcast data of the off-air anchor is determined; When the comparison result shows that the anchor information is reduced, the second final anchor information except the first final anchor information among all the second final anchor information is determined as the downcast anchor information, and the anchor corresponding to the downcast anchor information is determined as the downcast anchor; Determine the time node corresponding to the current frame as the target download time node, and select the target upload time node of the downloading anchor from the preset upload time record, wherein the upload time record includes anchor information of different anchors and the corresponding upload time nodes; The live broadcast data of the offline host is generated according to the first live broadcast data of the target live broadcast room at the target online broadcast time node and the second live broadcast data at the target offline broadcast time node.
[0064] Optional, such as Figure 4 As shown, the device also includes a time recording module 15, which is specifically used for: When the comparison result is a new addition to the anchor information, the time node corresponding to the current frame is determined as the broadcast time node; A mapping relationship is established between the newly added anchor information and the broadcast time node and added to the preset broadcast time record.
[0065] Optionally, the device further includes a data entry module 16, specifically used for: Obtaining the host information, host avatar and voice clip of at least one host to be recorded sent by the user's terminal; Based on a single anchor avatar, determine the facial sample identifier of the corresponding anchor avatar through a preset third-party face detection service; Based on a single voice segment, the voiceprint identifier of the corresponding voice segment is determined through a preset third-party intelligent voice recognition service; The anchor information, anchor avatar, face sample identification and voiceprint identification corresponding to each anchor to be recorded are stored in a preset database.
[0066] It should be noted that the above embodiment provides a kind of anchor live broadcast data generation device, when executing the anchor live broadcast data generation method, only takes the division of the above functional modules as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the above embodiment provides a kind of anchor live broadcast data generation device and a kind of anchor live broadcast data generation method embodiment, which belong to the same concept. The embodiment and implementation process thereof are detailed in the method embodiment, which will not be repeated here.
[0067] An embodiment of the present application further discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, a method for generating live broadcast data of an anchor in the above embodiment is adopted.
[0068] Among them, the computer program can be stored in a computer-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The computer-readable medium includes any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above-mentioned components.
[0069] Among them, through this computer-readable storage medium, a method for generating live broadcast data of an anchor in the above embodiment is stored in a computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.
[0070] An embodiment of the present application also discloses an electronic device, in which a computer program is stored in a computer-readable storage medium. When the computer program is loaded and executed by a processor, the above-mentioned method for generating live broadcast data of an anchor is adopted.
[0071] The electronic device may be a desktop computer, a laptop computer, a cloud server or other electronic device, and the electronic device includes but is not limited to a processor and a memory. For example, the electronic device may also include input and output devices, a network access device, and a bus.
[0072] Among them, the processor can adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.
[0073] Among them, the memory can be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device, or it can be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital card (SD) or a flash memory card (FC) equipped on the electronic device. Moreover, the memory can also be a combination of an internal storage unit and an external storage device of the electronic device. The memory is used to store computer programs and other programs and data required by the electronic device. The memory can also be used to temporarily store data that has been output or is to be output, and this application does not impose any restrictions on this.
[0074] Among them, through this electronic device, a method for generating live broadcast data of an anchor in the above-mentioned embodiment is stored in the memory of the electronic device, and is loaded and executed on the processor of the electronic device for easy use.
[0075] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. This application is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for generating live broadcast data of an anchor, characterized in that: The method comprises: Acquire a current frame of a real-time live broadcast screen in a target live broadcast room, wherein the current frame is a frame of the real-time live broadcast screen at a current time, and the current frame contains at least one anchor to be identified; Based on the live audio file corresponding to the current frame and the frame-cut interval, determine the first final anchor information of each anchor to be identified in the current frame, the frame-cut interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the last frame captured last time, and the duration between the current frame and the last frame is the periodic interval duration; Compare the first final host information corresponding to the current frame with the second final host information corresponding to each host in the previous frame to obtain a comparison result; According to the comparison result, the off-air host is determined and the live broadcast data of the off-air host is generated.
2. The method for generating anchor live broadcast data according to claim 1, characterized in that: The determining, based on the current frame picture and the live audio file corresponding to the frame-cutting interval, the first and final anchor information of each anchor to be identified in the current frame picture specifically includes: Based on the current frame, determine the first anchor information corresponding to each anchor to be identified; Determine the second host information of each host in the live audio file based on the live audio file corresponding to the frame clipping interval; Intersecting each of the first anchor information and each of the second anchor information, to obtain first final anchor information of each of the to-be-identified anchors in the current frame.
3. The method for generating anchor live broadcast data according to claim 2, characterized in that: The determining, based on the current frame, first anchor information corresponding to each anchor to be identified specifically includes: Based on the current frame, determining at least one face coordinate by a preset third-party face detection service; Based on the facial coordinates, the actual face of the corresponding host to be identified is captured from the current frame, and each actual face is compared with the host head portrait in a preset database to obtain a corresponding target face identification, wherein the database includes different host information and corresponding host head portraits and face sample identifications; The first anchor information corresponding to each target face identifier is matched from the database.
4. The method for generating anchor live broadcast data according to claim 2, characterized in that: The determining, based on the live audio file corresponding to the frame clipping interval, the second host information of each host in the live audio file specifically includes: Based on the live audio file corresponding to the frame-cut interval, at least one target voiceprint identifier is determined through a preset third-party intelligent voice recognition service; The target anchor information corresponding to each target voiceprint identifier is matched from a preset database, and the target anchor information is determined as the second anchor information of each anchor in the live audio file, wherein the database includes different anchor information and corresponding voiceprint identifiers.
5. The method for generating anchor live broadcast data according to claim 1, characterized in that: The step of determining the off-air host and generating the live broadcast data of the off-air host according to the comparison result specifically includes: When the comparison result is that the anchor information is newly added or the anchor information remains unchanged, the step of obtaining the current frame picture in the real-time live broadcast picture of the target live broadcast room is repeatedly performed after the interval time, until the comparison result is that the anchor information is reduced, and the live broadcast data of the anchor who is broadcasting is determined to be offline; When the comparison result is that the anchor information is reduced, the second final anchor information other than the first final anchor information among all the second final anchor information is determined as the downcast anchor information, and the anchor corresponding to the downcast anchor information is determined as the downcast anchor; Determine the time node corresponding to the current frame as the target download time node, and select the target upload time node of the downloading anchor from the preset upload time record, wherein the upload time record includes anchor information of different anchors and the corresponding upload time nodes; The live broadcast data of the offline host is generated according to the first live broadcast data of the target live broadcast room at the target broadcast time node and the second live broadcast data of the target offline broadcast time node.
6. The method for generating anchor live broadcast data according to claim 5, characterized in that: The method further comprises: When the comparison result is a new addition to the broadcast information, the time node corresponding to the current frame is determined as the broadcast time node; A mapping relationship is established between the newly added anchor information and the broadcast time node and added to the preset broadcast time record.
7. The method for generating anchor live broadcast data according to claim 3 or 4, characterized in that: The method further comprises: Obtaining the host information, host avatar and voice clip of at least one host to be recorded sent by the user's terminal; Based on the single anchor avatar, determine the face sample identifier of the corresponding anchor avatar through a preset third-party face detection service; Based on the single voice segment, determine the voiceprint identifier of the corresponding voice segment through a preset third-party intelligent voice recognition service; The anchor information, anchor avatar, face sample identification and voiceprint identification corresponding to each anchor to be recorded are stored in a preset database.
8. A device for generating live broadcast data, characterized in that: include: An information acquisition module (11) is used to acquire a current frame of a real-time live broadcast picture of a target live broadcast room, wherein the current frame is a picture frame of the real-time live broadcast picture at a current time, and the current frame contains at least one anchor to be identified; The anchor identification module (12) is used to determine the first and final anchor information of each anchor to be identified in the current frame based on the live broadcast audio file corresponding to the current frame and the frame-cut interval, wherein the frame-cut interval is the time interval between the time node corresponding to the current frame and the time node corresponding to the last frame captured last time, and the time length between the current frame and the last frame is the periodic interval time length; An information comparison module (13) is used to compare the first final anchor information corresponding to the current frame with the second final anchor information corresponding to each anchor in the previous frame to obtain a comparison result; A data generation module (14) is used to determine the off-air host and generate live broadcast data of the off-air host according to the comparison result.
9. A computer-readable storage medium having a computer program stored therein, characterized in that: When the computer program is loaded and executed by a processor, the method according to any one of claims 1 to 7 is adopted.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor loads and executes the computer program, the method according to any one of claims 1 to 7 is adopted.