Multi-person identity recognition method and device in live broadcast stream, equipment and medium

By constructing the distance matrix and solving the optimal matching, the problem of low facial recognition accuracy in live videos is solved, the accuracy and stability of multi-person identity recognition are achieved, and the user interaction experience is enhanced.

CN119996718APending Publication Date: 2025-05-13GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510107713.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art faces the problems of low facial recognition accuracy, inability to process low-quality facial images, and inability to achieve global optimal matching in live video scenarios, resulting in poor identity recognition errors and interactive experience in multi-person live broadcasts.

Method used

By constructing a distance matrix and solving the optimal match, the unique matching relationship between each face image and the face template is determined, the identity identification is accurately recognized, and the user's operation instructions are responded to the mapping relationship data.

Benefits of technology

It improves the accuracy and stability of multi-person identity identification in live streams, reduces identity mismatch and duplicate matching problems, and enhances user interaction experience and the interactiveness of live streaming platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996718A_ABST
    Figure CN119996718A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network live broadcast, and discloses a multi-person identity recognition method and device in a live broadcast stream, equipment and a medium, and the method comprises the steps: determining a data distance between each face image in the live broadcast stream of a live broadcast room and a plurality of face templates corresponding to a plurality of anchor persons in a preset live broadcast person library corresponding to the live broadcast room, obtaining a distance matrix; the optimal matching between the face images and the face templates is solved according to the distance matrix, and the identity label of the identity of the anchor person to which the uniquely matched face template of each face image belongs is determined through the optimal solution; and responding to an operation instruction acting on the identity of the target anchor person according to the position information of each face image in the video frame and the mapping relation data between the identity label of the face template matched with the face image. According to the method, the anchor person identity is determined by searching the optimal matching, so that the accuracy of multi-person identity recognition in the live broadcast stream and the user experience are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to network live broadcast technology, and in particular to a method for identifying multiple people in a live broadcast stream and its device, equipment, and medium. Background Art

[0002] With the development of network technology and the rise of social media, live streaming has become an indispensable part of people's daily lives. In the field of live streaming, especially in the scenario where multiple anchors are broadcasting live through the same camera at the same time, users not only expect to be able to watch the live content, but also hope to have a deeper interaction with any anchor. In order to achieve these functions, it is necessary to be able to identify each anchor in the live video so as to correctly respond to user instructions acting on the anchor.

[0003] However, existing face recognition technology faces many challenges in live video scenarios. First, the dynamic nature of live video causes the face area in the picture to often have problems such as side faces, lowered heads, occlusions, exaggerated expressions, and motion blur. These problems make it difficult for traditional face recognition algorithms to accurately identify and match faces. Secondly, multi-person scenes in live broadcasts increase the complexity of recognition, because the faces of different anchors may overlap or intersect in the picture, making it difficult for the recognition algorithm to distinguish and match. In addition, different anchors may require different beauty and makeup parameters, and traditional technologies cannot accurately apply these parameters to the correct anchor's face.

[0004] In traditional technologies, face recognition usually requires high-quality face images, such as frontal, clear, and unobstructed photos. However, in the live broadcast environment, it is difficult to obtain face images that meet these conditions due to the above reasons. Therefore, when processing low-quality face images, existing algorithm models often cannot achieve accurate face recognition, resulting in problems such as mismatch and duplication of person identity tags. These problems not only affect the interactive experience of live broadcasts, but also limit the design and innovation of live broadcast gameplay.

[0005] Traditional technologies often use a simple nearest neighbor matching method when dealing with multi-person live broadcasts, that is, calculating the similarity between each detected face and all feature vectors in the template library, and then selecting the template with the highest similarity as the matching result. This method is prone to errors when the face image quality is low, because it does not take into account the overall optimal match, but only maximizes the single matching result. This results in multiple different faces being mistakenly matched to the same template in multi-person live broadcast scenarios, or multiple expressions of the same person being matched to different templates, making it impossible to achieve accurate one-to-one recognition.

[0006] In summary, traditional technologies have problems in multi-person identity recognition in live streaming, such as low recognition accuracy, inability to process low-quality face images, and inability to achieve global optimal matching. These problems seriously limit the improvement of live interactive experience and the further development of live technology. Therefore, it is necessary to develop a new technical solution to solve the problems existing in traditional technologies. Summary of the invention

[0007] The purpose of the present application is to solve the above-mentioned problem and to provide a method for identifying multiple people in a live stream and its corresponding device, equipment, and non-volatile readable storage medium.

[0008] According to one aspect of the present application, a method for identifying multiple people in a live stream is provided, comprising the following steps: determining the data distance between each facial image in the live stream of a live broadcast room and a plurality of facial templates corresponding to a plurality of anchor characters in a preset live broadcast character library corresponding to the live broadcast room, and obtaining a distance matrix; solving the optimal match between the facial image and the facial template based on the distance matrix, so that the sum of the data distances between the facial image and the facial template is the optimal solution, so as to determine the identity identifier of the anchor character identity to which the facial template uniquely matches each facial image belongs; responding to an operation instruction acting on the target anchor character identity according to the mapping relationship data between the position information of each facial image in a video frame and the identity identifier of the facial template matching the facial image.

[0009] According to another aspect of the present application, a device for identifying multiple people in a live stream is provided, including: a similarity operation module, configured to determine the data distance between each face image in the live stream of a live broadcast room and a plurality of face templates corresponding to a plurality of anchor characters in a preset live broadcast character library corresponding to the live broadcast room, to obtain a distance matrix; a face matching module, configured to solve the optimal match between the face image and the face template based on the distance matrix, so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity identifier of the anchor character identity to which the face template uniquely matching each face image belongs; a mapping response module, configured to respond to an operation instruction acting on the target anchor character identity according to the mapping relationship data between the position information of each face image in a video frame and the identity identifier of the face template matching the face image.

[0010] According to another aspect of the present application, a device for identifying multiple people in a live stream is provided, comprising a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the method for identifying multiple people in a live stream described in the present application.

[0011] According to another aspect of the present application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the method for identifying multiple people in a live stream in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the method are executed.

[0012] This application effectively addresses the challenges in the prior art and achieves significant beneficial effects: by constructing a distance matrix and solving the optimal match to process facial images in live videos, the recognition accuracy is improved, and it can cope with a variety of complex scenarios and reduce the mismatch and duplication of anchor identities; at the same time, it achieves global optimal matching based on the distance matrix, ensuring the stability and reliability of the matching results; in addition, by mapping relational data, it can accurately respond to users' operating instructions to specific anchors, such as sending virtual gifts or applying beauty effects, greatly improving the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is an exemplary network architecture suitable for applying the method for identifying multiple people in a live stream of the present application;

[0014] Figure 2 A flowchart of an embodiment of a method for identifying multiple people in a live stream of the present application;

[0015] Figure 3 This is a principle block diagram of a device for identifying multiple people in a live stream of the present application;

[0016] Figure 4 This is a schematic diagram of the structure of a multi-person identity recognition device in a live stream used in this application. DETAILED DESCRIPTION

[0017] Before introducing the specific embodiments of the technical solution of the present application in detail, a network architecture and application scenarios suitable for supporting the implementation of the technical solution of the present application are first disclosed.

[0018] The technical solution of the present application is applicable to the field of live broadcasting, and is particularly applicable to scenarios where multiple identities need to be identified in a live broadcast stream. In this context, the technical solution of the present application can be applied in a typical live broadcasting platform, such as Figure 1 As shown, the platform is composed of a live broadcast server 83, a media server 85, a terminal device 90 of the anchor user and a terminal device 92 of the audience user.

[0019] The live broadcast server 83 is responsible for managing the creation of the live broadcast room and the reception and distribution of the live broadcast stream. As the central node of the live broadcast stream, it receives the live broadcast stream from the anchor user terminal device 90 and distributes it to the media server 85 and the audience user terminal device 92. The live broadcast server 83 is also responsible for handling the management tasks of the live broadcast room, such as user authentication, live broadcast room settings, etc.

[0020] The media server 85 is responsible for encoding, transcoding and storing the live stream. It performs necessary transcoding processing on the live stream according to the network conditions and device capabilities of the audience users to ensure that the live stream can be transmitted to the audience in an appropriate format and quality. The media server 85 is also responsible for storing the live stream for playback or other subsequent processing.

[0021] The anchor user's terminal device 90 is the source of the live content. The anchor captures the video through the terminal device 90 and sends the original video stream to the live server 83. The anchor user's terminal device 90 can be a smart phone, tablet computer, laptop computer or professional camera equipment, which is connected to the live server 83 via the Internet.

[0022] The terminal device 92 of the audience user is the receiving end of the live content. The audience watches the live broadcast through the terminal device 92, which can be a smart phone, a tablet computer, a personal computer or a smart TV. The terminal device 92 of the audience user is connected to the live broadcast server 83 through the Internet, receives the live broadcast stream and plays it.

[0023] The transmission process of the live stream is as follows: the anchor user's terminal device 90 captures the video and notifies the live server 83, which distributes the live stream to the media server 85. The media server 85 encodes and transcodes the live stream, and then sends the network address of the processed live stream back to the live server 83, which then transmits the network address of the live stream to the viewer user's terminal device 92, so that the terminal device 92 can load and play the live stream according to the network address.

[0024] Programming according to the method for identifying multiple people in a live stream of the present application can be implemented as a computer program product, which can be selectively deployed in the terminal device 90 of the anchor user, the media server 85 or the terminal device 92 of the audience user. The program product contains the algorithm and logic required to implement the technical solution of the present application, and can identify the identities of multiple anchor characters in the video frame of the live stream, so as to carry out a variety of subsequent applications, including but not limited to beautification processing corresponding to beauty and makeup, and giving virtual gifts to specific anchor characters.

[0025] See also Figure 2 According to a method for identifying multiple people in a live stream provided by the present application, the method can be implemented as a computer program product, which is installed and run in various node devices of a live broadcast platform, such as a terminal device of a host user, a media server or a terminal device of an audience user, to be responsible for implementing the identity identification processing of multiple host characters in a video frame in a live stream. In some embodiments of the method, the following steps are included:

[0026] Step S3100, determining the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room, and obtaining a distance matrix;

[0027] When a host user invites multiple hosts to participate in a live broadcast event, the host user's terminal device camera collects live video images, forms a live stream and transmits it to the live broadcast room for live broadcast. After the audience user's terminal device obtains this live stream, it decodes and plays it, so that the audience can see the display image of the live stream in the graphical user interface, thereby recognizing each host character. As the host users of the online live broadcast platform, these hosts have their own user IDs, which are their identity identifiers. At the computer program level, each host character can also be recognized through the identity identifier.

[0028] In order to accurately identify the identity of each anchor character in the live stream, the anchor user needs to prepare a live character library in advance. This library stores the facial images of each anchor character participating in the live broadcast, and these images are used as face templates for subsequent identity matching and recognition. In one embodiment, the face template can be directly uploaded to the live character library by the anchor user responsible for hosting the multi-person live broadcast event. In another embodiment, each anchor character participating in the multi-person live broadcast can provide his or her face image as a face template in association with his or her account. When the anchor user hosting the multi-person live broadcast event configures his or her live character library, he or she can select one or more anchor characters participating in the live broadcast from the list of friends he or she follows. In response to the selected event of the anchor user in the current live broadcast room, according to the identity identifier of the selected anchor character, the face template of the anchor character is obtained from the corresponding anchor character account and stored in the live character library, which is more convenient and efficient.

[0029] For each face template in the live broadcast character library, the image feature information of the face template can be extracted with the help of a pre-trained feature extraction model to achieve feature representation of each face template. This model can identify and extract key semantic features in face images and represent them as high-dimensional vectors such as 512 or 1024 dimensions. These high-dimensional vectors contain feature information of face images and are used for subsequent distance calculation and matching. The host character's identity and the high-dimensional vector of its corresponding face template are associated with the face template and stored in the live broadcast character library, providing a basis for data distance calculation between face images in the live stream and preset face templates.

[0030] In specific implementation, the feature extraction model can use a deep learning-based convolutional neural network (CNN) for classification training to learn the ability to extract effective high-dimensional feature vectors from face images. These feature vectors not only contain the geometric information of the face, but also contain detailed information such as texture and expression, so that each face template can be uniquely identified. In this way, the live character library can provide an accurate reference for each face image in the live stream, so as to perform real-time identity recognition and matching during the live broadcast process.

[0031] During the live broadcast process, the video frames of the live broadcast stream contain moving images of multiple anchor characters. First, it is necessary to extract the facial images of each anchor character from these continuous video frames. Accordingly, various mature human body detection technologies can be used to first determine the human body image, and then use face detection technology to determine the facial image based on the human body image, and finally locate and segment the facial area from the video frame. Specifically, the terminal device or media server using the method of the present application will receive the live broadcast stream in real time, and use the preset human body and face detection algorithm to process each frame of the image, identify the face therein and extract it, and form a real-time face set. These detected face images are then used to compare with the face templates in the live broadcast character library.

[0032] The extracted face image then needs to be converted into a format that can be used for calculation, namely a high-dimensional vector. This conversion process is completed through the feature extraction model revealed in the previous article, which can extract key semantic features from the face image and encode them into vectors in a high-dimensional space, namely high-dimensional vectors. These high-dimensional vectors capture the key features of the face, such as the location of the eyes, nose and mouth, as well as facial expressions and contours.

[0033] After having the high-dimensional vector representation of the face image, the next step is to calculate the data distance between these vectors and each face template in the live character library. The calculation of data distance usually uses measurement methods such as cosine similarity, Euclidean distance, Minkowski distance, Jaccard similarity coefficient, Manhattan distance, Chebyshev distance, and Pearson correlation coefficient. Cosine similarity evaluates the similarity of two vectors by measuring the angle between them, while Euclidean distance measures the straight-line distance between two points in Euclidean space. These calculation results form a distance matrix, in which each row represents a detected face image, each column represents a face template in the live character library, and each element in the matrix is ​​the distance value between the corresponding face image and the face template.

[0034] In the above data distance algorithms, the distance values ​​have different directions of representation for the degree of similarity. The larger the values ​​of cosine similarity, Jaccard similarity coefficient, Pearson correlation coefficient and Spearman rank correlation coefficient, the more similar the face image and face template are; while the smaller the values ​​of Euclidean distance, Manhattan distance and Chebyshev distance, the more similar the face image and face template are. In practical applications, this representation can also be converted to each other as needed. For example, by simply adding a negative sign to the distance value of the Euclidean distance, it can be converted into a representation in which the larger the value, the more similar the face image and face template are.

[0035] Step S3200, solving the optimal match between the face image and the face template based on the distance matrix, so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity of the anchor character to which the face template uniquely matches each face image belongs;

[0036] In the process of multi-person identification in live streaming, ensure that each face image identified from the video frame can be matched with a unique face template, and make the sum of the data distances of all matching pairs reach the optimal solution. The definition of the optimal solution depends on the direction represented by the sum of data distances: if the larger the distance value, the higher the similarity, then the optimal solution is the solution with the largest sum of data distances; if the smaller the distance value, the higher the similarity, then the optimal solution is the solution with the smallest sum of data distances.

[0037] In order to achieve the best match, a variety of algorithms can be used, one of which is an algorithm for finding the minimum match, which can find the optimal solution in multiple choice problems. This algorithm constructs a cost matrix and then finds the matching solution with the minimum cost through row and column operations. In this application, the cost matrix is ​​a distance matrix composed of the data distance between the face image and the face template. Since this algorithm is looking for the best match, when the direction represented by the distance value in the distance matrix is ​​that the larger the distance value, the higher the similarity, it should be converted, such as adding a negative sign to negate it, so that the smaller the distance value, the higher the similarity. The algorithm iteratively updates the matching state by finding an augmented path until the best match is found.

[0038] Another algorithm is the linear programming method, which solves the optimal matching problem by constructing a linear objective function and a set of linear constraints. The objective function aims to minimize or maximize the sum of data distances, while the constraints ensure that each face image is matched with only one face template. The linear programming problem can be solved by algorithms such as the simplex method or the interior point method.

[0039] Greedy algorithm is also a feasible method, which takes the best choice in the current state in each step of selection in the hope of leading to the global optimal solution. In this application, the greedy algorithm can be used to select the minimum data distance for matching in each round until all face images are matched.

[0040] In addition to the above algorithms, you can also use matching algorithms based on deep learning. This type of algorithm trains a deep neural network to learn the complex matching relationship between face images and face templates to achieve optimal matching. Deep learning models can automatically extract features and perform matching, reducing the need for manual feature engineering.

[0041] In practical applications, the choice of algorithm depends on the nature and scale of the specific problem, as well as the available computing resources. For example, for smaller problems, algorithms for finding the minimum match and linear programming methods may be more appropriate because they can guarantee the global optimal solution. For large-scale problems, greedy algorithms or deep learning-based algorithms may be more practical because they are more computationally efficient. Through the application of these algorithms, the optimal match for multi-person identity recognition in live streams can be achieved, improving the accuracy and efficiency of recognition.

[0042] After the matching relationship between the face image and the face template is determined, since the face template and the identity identifier of the anchor character to which it belongs have been associated in advance, the relationship between the face image and the identity identifier of the anchor character to which it belongs is also determined.

[0043] Step S3300: respond to the operation instruction acting on the identity of the target anchor character according to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image.

[0044] After determining the matching relationship between each face image and the corresponding face template, the position information of the face image in the video frame can be mapped to its corresponding identity to form a mapping table. This mapping table is the key to achieving accurate operation command response, which allows various operation commands for specific anchor characters to be executed based on the position and identity of the face image.

[0045] For example, in one embodiment, when a viewer user clicks on a face image of a certain anchor character in a graphical user interface, the anchor character identity corresponding to the clicked face image can be quickly determined through a mapping table based on the viewer user's click position, thereby executing the operation the viewer user wants to perform, such as sending a virtual gift, initiating an interaction, etc., allowing the user to have a convenient and fast experience.

[0046] In another embodiment, the audience user or the anchor user directly specifies the identity of the anchor character through an operation instruction. In this case, the identity can be directly used to query the mapping table to determine the location information of the corresponding face image, and the corresponding image processing effect, such as beauty filter or dynamic sticker, is synthesized into the face image of the anchor character in the video frame according to the location information. The advantage of this method is that it does not rely on the user's click operation on the face image, but is directly associated with a specific anchor character through the identity, making the operation more flexible and direct.

[0047] In another embodiment, the audience user clicks on the face image of a certain anchor character in the video, and the mapping table is queried according to the position clicked by the audience user to determine the designated face image and its identity identifier, and then the audience user enters the comment text to trigger the operation instruction. In response to the operation instruction, the comment text can be synthesized as a bullet screen around the face image of the anchor character, and a corresponding directional mark can be added. This function not only enhances the interactivity of the live broadcast, but also provides a more personalized viewing experience for the audience, making the direction of the bullet screen clearer.

[0048] The implementation of these operation instructions depends on the establishment of accurate identity recognition and mapping relationship data. In this way, it can ensure that the operation instructions are accurately applied to the target anchor, whether it is through image clicks or identity identification. This precise response mechanism improves the interactivity of live broadcasts and the participation of viewers, while also providing more business opportunities for live broadcast platforms, such as accurate advertising and personalized content recommendations associated with specific anchors.

[0049] Through the application of the above embodiments, the online live broadcast platform can realize accurate identity recognition of multiple anchor characters in the live broadcast video, thereby significantly improving the interactivity and user experience of the live broadcast. Its technical advantages include but are not limited to:

[0050] First, by building a live broadcast character database and real-time face detection technology, we ensure that in a multi-person live broadcast scenario, each person's face image can be quickly and accurately detected and extracted. This process not only improves recognition efficiency, but also provides accurate basic data for subsequent identity matching.

[0051] Secondly, by using the facial feature vector extracted by the deep learning model, each facial image can be uniquely identified, which greatly enhances the accuracy of recognition. Compared with traditional face recognition technology, this method can handle low-quality facial images, such as side faces, lowered heads, and occlusions, which is particularly important in a dynamic live broadcast environment. In this way, even if the position and expression of the host character changes during the live broadcast, high accuracy recognition can be maintained.

[0052] Furthermore, by solving the optimal matching problem based on the distance matrix between the face image and the face template, the best match between the face image and the face template is achieved, and the sum of the data distance is optimized, thereby ensuring that each face image can be matched with the correct face template. This global optimization matching strategy avoids the problems caused by local optimal solutions, such as identity mismatch or repeated matching, and improves the stability and reliability of matching.

[0053] In addition, by establishing a mapping relationship between the location information of the face image and the identity identifier, it is possible to respond to various operation instructions for specific anchor characters. Whether it is a click operation by the audience user or an instruction input based on the identity identifier, the target anchor character can be quickly and accurately located and the corresponding operation can be performed, such as sending virtual gifts, applying beauty effects, or displaying bullet comments. This response mechanism not only enhances the audience's sense of participation, but also creates more interactive opportunities and commercial value for the live broadcast platform.

[0054] On the basis of any embodiment of the method of the present application, the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room is determined to obtain a distance matrix, including:

[0055] Step S3111: extracting video frames from the live stream transmitted in real time in the live broadcast room, and performing face tracking detection on the video frames to obtain multiple face images therein to form a real-time face set;

[0056] In the process of real-time transmission of live streaming, the live streaming is a continuous video signal collected and transmitted by the camera of the anchor user's terminal device. In this regard, video frame images can be extracted frame by frame from the live streaming, and face tracking detection can be performed on each frame, with the purpose of identifying and locating the facial areas of all anchor characters appearing in the frame, intercepting the images therein, and obtaining the corresponding face images to form a real-time face set.

[0057] The method of obtaining the video frame of the live stream varies according to the different devices deployed by the method of the present application. For example, for the terminal device of the anchor user, the video frame can be directly obtained from the image data collected by the camera, while at the terminal device of the media server and the audience user, since it will first receive the live stream for decoding, the corresponding video frame can be extracted from the decoded image data.

[0058] Face tracking detection technology usually relies on advanced computer vision algorithms that can accurately identify the location and features of faces in video frames. In one embodiment, a pre-trained face detection model can be used to implement it. The model can handle various complex situations, including changes in the angle of the face, changes in expression, and changes in lighting conditions. By applying the face detection model, multiple face images can be extracted from each video frame and these images can be assembled to form a real-time face set.

[0059] In specific implementation, a variety of face detection technologies can be used, such as cascade classifiers based on Haar features, multi-task cascade convolutional neural networks (MT-CNN) based on deep learning, or single shot detector enhancement (SSD). These methods have their own advantages, such as the advantages of cascade classifiers in processing speed, or the advantages of deep learning methods in accuracy. Choosing the right method depends on the specific application scenario and performance requirements.

[0060] Step S3112, calculating the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and expressing the data distance relationship between each face image and each face template as a distance matrix;

[0061] Once the real-time face set is obtained, the data distance between these real-time face images and the preset face templates in the live broadcast character library can be calculated. In order to achieve this calculation, it is first necessary to extract the feature vector of each face image from the real-time face set. These feature vectors are high-dimensional representations automatically learned from face images through deep learning models, such as convolutional neural networks (CNN). These vectors can capture the key features of the face, including geometric structure, texture, and expression. Similarly, each face template in the live broadcast character library has also been converted into a corresponding feature vector and stored in association with the identity of the anchor character.

[0062] When calculating data distance, you can use a variety of metrics, such as cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, etc. Cosine similarity measures the similarity of two vectors by calculating the angle between them, which is suitable for scenarios where direction rather than size is of concern. Euclidean distance measures the straight-line distance between two points in Euclidean space, which is suitable for scenarios where absolute differences are of concern. Manhattan distance and Chebyshev distance are also commonly used metrics, which respectively calculate the walking distance and the difference in the maximum dimension between two points in the urban block model.

[0063] In specific implementation, one or more measurement methods can be selected to calculate the distance between the real-time face image and each face template. For example, if the direction information of the face image is of interest, cosine similarity can be selected; if the absolute difference of the face image is of interest, Euclidean distance can be selected. These distance values ​​will be organized into a distance matrix, where each row represents a real-time face image, each column represents a preset face template, and each element in the matrix represents the distance value between the corresponding face image and the face template.

[0064] The construction of the distance matrix provides a basis for the subsequent optimal matching solution. By analyzing this matrix, it is possible to determine which face images are most similar to a specific face template, thereby achieving accurate anchor character identification. This process not only improves the accuracy of identification, but also provides technical support for subsequent interactive operations, such as giving virtual gifts and applying beauty and makeup effects.

[0065] Step S3113: Detect whether the number of rows and columns of the distance matrix are the same. If they are different, add the missing rows or columns and assign zero values ​​to each element therein to update the distance matrix.

[0066] The number of rows in the distance matrix represents the number of face images in the real-time face set, and the number of columns represents the number of face templates in the live character library. In order to perform effective matching calculations, in this embodiment, the distance matrix is ​​standardized as a square matrix, that is, the number of rows and columns must be the same. If the number of rows and columns is inconsistent, the matrix needs to be supplemented to ensure that the number of rows and columns of the matrix is ​​equal.

[0067] Specifically, if the number of face images in the real-time face set is greater than the face templates in the live character library, or vice versa, the distance matrix needs to be adjusted. This can be done by adding the missing rows or columns and assigning zero values ​​to each element in these newly added rows or columns. For example, if the distance matrix originally has 5 rows and 6 columns, it means that there are only 5 face images in the real-time face set, and there are 6 face templates in the live character library, then you need to add a row to the matrix and set all elements of this row to zero, so that the matrix becomes 6 rows and 6 columns. Similarly, if there is one more face image in the real-time face set relative to the live character library, you also need to add corresponding columns to the matrix and set all elements of these columns to zero.

[0068] This supplementation operation ensures that the distance matrix is ​​suitable for the subsequent optimal matching algorithm, which requires that the distance matrix must be a square matrix. In this way, even if the number of face images and face templates is inconsistent, effective matching calculations can be performed. The supplemented matrix provides the basis for distance value adjustment and matching in subsequent steps, allowing the algorithm to find the optimal match and thus determine the identity of the anchor person corresponding to each face image.

[0069] It should be pointed out that, since the distance values ​​in the distance matrix, i.e., the numerical values ​​of its elements, are different in the direction of the similarity between the face image and the face template, depending on the data distance algorithm applied. To facilitate the implementation of the matching algorithm of this embodiment, the distance values ​​of this distance matrix can be converted into a numerical form in which the smaller the distance value, the more similar the face image and the face template are. To this end, when the larger the distance value, the more similar the face image and the face template are, the distance value can be added with a negative sign to negate it. In this case, the optimal matching algorithm of this embodiment can be implemented according to the minimum matching algorithm.

[0070] This embodiment provides a standardized working environment for the subsequent optimal matching algorithm by ensuring that the number of rows and columns of the distance matrix is ​​the same, thereby bringing significant technical advantages.

[0071] First, this method solves the problem of inconsistency between the number of face images and the number of face templates by filling in the missing rows or columns and assigning zero values, enabling the algorithm to perform matching calculations on a unified basis, thereby improving the adaptability and flexibility of the matching process.

[0072] Secondly, the normalized distance matrix provides a clear framework for the algorithm, which enables it to accurately find the optimal match between each face image and the corresponding face template even in a complex and changing live broadcast environment, thereby improving the accuracy and efficiency of recognition.

[0073] In addition, by converting the distance matrix into a form in which the smaller the distance value, the higher the similarity, this embodiment simplifies the implementation of the matching algorithm, allowing the algorithm to focus more on finding a solution that minimizes the total distance. This conversion not only improves the execution efficiency of the algorithm, but also makes the algorithm easier to implement and maintain. Ultimately, this method enables the live broadcast platform to quickly respond to user operation instructions during real-time live broadcasts, such as sending virtual gifts or applying beauty effects, greatly improving the user's interactive experience and satisfaction. In this way, this embodiment not only improves the accuracy and efficiency of multi-person identity recognition in live broadcast streams, but also enhances the interactivity and viewing experience of the live broadcast platform, creating more business opportunities and user value for the live broadcast platform.

[0074] On the basis of any embodiment of the method of the present application, the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room is determined to obtain a distance matrix, including:

[0075] Step S3131, extracting video frames from the live stream transmitted in real time in the live broadcast room, performing face tracking detection on the video frames to obtain multiple face images therein to form a real-time face set;

[0076] The implementation of this step is the same as step S3111 and will not be repeated here.

[0077] Step S3132, calculating the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and expressing the data distance relationship between each face image and each face template as a distance matrix;

[0078] The implementation of this step is the same as step S3112 and will not be repeated here.

[0079] Step S3133, calculate the larger number of rows and columns of the distance matrix, and for each row / column with a larger number, calculate the mean distance corresponding to the element values ​​of each column / row, and delete the rows / columns whose distance means represent the lowest similarity, so that the number of rows and columns of the distance matrix is ​​consistent.

[0080] In the live stream, the number of people in the real-time face set may be inconsistent with the number of face templates in the live character library for various reasons. To solve this problem, the distance matrix needs to be adjusted so that the number of rows and columns is equal, so as to ensure that each face image can correspond to a face template.

[0081] Specifically, if the number of rows in the distance matrix (representing the number of people in the real-time face set) is found to be greater than the number of columns (representing the number of face templates in the live character library), then those redundant rows need to be identified and deleted. Conversely, if the number of columns is greater than the number of rows, then the redundant columns need to be deleted.

[0082] In order to determine which rows or columns should be deleted, the mean of the element values ​​of each row and column in the distance matrix can be calculated. In the case where the number of rows is greater than the number of columns, the mean of each row is calculated, and then several rows with the largest mean are selected for deletion. The row with the largest mean means that the distance between the face image corresponding to the row and all face templates is large, indicating that the similarity between these images and the templates in the library is low, so these rows can be considered redundant. In the case where the number of columns is greater than the number of rows, the mean of each column is also calculated, and several columns with the largest mean are selected for deletion.

[0083] For example, if the distance matrix has 6 columns and 7 rows, it means that there are 7 face images in the real-time face set, but there are only 6 face templates in the live character library. At this time, calculate the mean of each row and find the rows with the largest mean. The face images corresponding to these rows have the lowest similarity with the templates in the library, so they can be deleted from the distance matrix until the number of rows and columns is equal.

[0084] This embodiment achieves significant technical advantages in the field of multi-person identity recognition in live streaming by further applying step S3133 based on the distance matrix, including but not limited to:

[0085] First, by adjusting the number of rows and columns of the distance matrix, we ensure that the number of people in the real-time face set matches the number of face templates in the live character library, thus avoiding matching errors caused by inconsistent numbers. This method improves the accuracy of matching because it ensures that each face image can find a corresponding face template, reducing recognition errors caused by insufficient or redundant templates.

[0086] Secondly, by calculating the mean of each row and column in the distance matrix, those face images with low similarity to the templates in the library are intelligently identified and deleted. This process not only optimizes the matching process, but also improves recognition efficiency because it reduces unnecessary calculations. Especially when processing large-scale data, this method can significantly reduce computational complexity and resource consumption.

[0087] Furthermore, this embodiment also enhances adaptability and flexibility. In a live broadcast environment, changes in the number of people are common, and step S3133 provides a dynamic adjustment mechanism that can quickly respond to changes in the number of people and always maintain efficient matching performance. This mechanism makes the identity recognition process more stable and ensures the continuity and consistency of the recognition results even when the number of people fluctuates.

[0088] In addition, the application of step S3133 also improves the interactive experience of live broadcasting. By ensuring that each face image can be accurately matched to the corresponding anchor character, it can more accurately respond to the audience's operation instructions, such as giving virtual gifts and applying beauty effects. This not only enhances the audience's sense of participation, but also creates more business opportunities for the live broadcast platform, such as accurate advertising and personalized content recommendations.

[0089] On the basis of any embodiment of the method of the present application, according to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image, responding to the operation instruction acting on the identity of the target anchor character includes:

[0090] Step S3311, responding to a touch operation instruction acting on the target anchor character identity in the display image of the video frame in the graphical user interface, determining the identity identifier of the target anchor character identity according to mapping relationship data between position information of each face image in the video frame and the identity identifier of a face template matching the face image;

[0091] In a live broadcast room, audience users often interact with video content through a graphical user interface. The present application adapts to this demand and provides an interactive method so that audience users can click on the target anchor character in the display screen of the video frame to trigger corresponding touch operation instructions for human-computer interaction.

[0092] In order to respond to such touch operation instructions, it is necessary to be able to identify the identity of the host being clicked. Therefore, the position information of each face image in the video frame can be determined in advance. When the face image is identified using face detection technology, the corresponding selection box of the face image in the video frame has been determined by the corresponding model, and it is determined in advance based on this selection box. These face images have also been matched with the face templates in the preset live broadcast character library, thereby establishing the mapping relationship data between the position information and the identity identifier.

[0093] When the audience performs a touch operation on the graphical user interface, such as clicking on a certain anchor character, by locating the clicked position, the corresponding face image can be retrieved through a mapping table that reflects the mapping relationship data between the position information of the face image in the video frame and the identity identifier of the face template that matches the face image. By using the previously established mapping table to query the position information, the face template corresponding to the face image can be quickly determined, and then the identity identifier of the target anchor character can be identified.

[0094] For example, in a live broadcast scene, if the audience user clicks on a host character in the video frame, the location information covering the coordinate position can be determined by searching the mapping table according to the clicked coordinate position. According to the mapping relationship between this location information and the identity identifier of the face template, the identity identifier of the host character corresponding to the clicked face image can be determined. It should be pointed out that this process can also be applied to touch operations for a variety of different purposes, such as likes, comments, or sending emoticons, etc., providing users with a way to directly interact with the live content.

[0095] It can be seen that by querying the mapping table through the coordinate position clicked by the audience user to determine the target anchor character targeted by the audience user, it is possible to effectively respond to the audience's touch operation instructions for the identity of the target anchor character in the video frame, realize the rapid matching between the face image and the identity identification in the live stream, and enhance the interactivity and user experience of the live broadcast.

[0096] Step S3312: a virtual gift selection interface is popped up to obtain gift configuration information, which includes the virtual gift determined by the user to be given to the target anchor character;

[0097] During the live broadcast interaction, audience users often express their support and love for the anchor by giving virtual gifts. In order to realize this interactive function, it is necessary to provide a mechanism that allows users to select and give virtual gifts to the target anchor character they choose. When the audience user clicks the target anchor character on the graphical user interface, a virtual gift selection interface can pop up in response to this operation instruction. This interface is used to obtain the user's gift configuration information, including the virtual gifts that the user chooses to give to the target anchor character.

[0098] The virtual gift selection interface can display a series of virtual gift options, each of which can be equipped with corresponding animated images and descriptions. Users can select one or more virtual gifts from these options based on their personal preferences and the emotions they want to express. After making their selections, the user submits the gift-giving operation instruction, and the background will construct the gift-giving configuration information of the audience user and perform subsequent processing.

[0099] After the audience user determines the virtual gift based on the virtual gift selection interface and triggers the background to determine the gift configuration information, it submits it to the live broadcast server, and the live broadcast server can control the media server to perform subsequent steps according to the gift configuration information.

[0100] Step S3313, responding to the gift giving operation instruction submitted by the user, obtaining the animated image of the virtual gift in the gift giving configuration information, and synthesizing the animated image into the face image of the target anchor character identity corresponding to the identity identifier in each video frame of the live stream.

[0101] When the viewer user selects and submits a virtual gift on the graphical user interface, the live broadcast server receives the gift-giving operation instruction and obtains the corresponding gift-giving configuration information. This information contains the details of the virtual gift selected by the user, including the animated image of the virtual gift and other related settings.

[0102] The live broadcast server then sends these configuration information to the media server, instructing the media server to perform corresponding image processing operations according to the gift configuration information. After receiving the instructions from the live broadcast server, the media server will determine the animated images of the virtual gifts from its stored resources or real-time content. These animated images are pre-designed, corresponding to the virtual gifts, and stored on the server in a computer-readable form.

[0103] Next, the media server decodes the live stream of the anchor user it receives, performs image processing on the decoded live stream video frame, and synthesizes the animated image of the virtual gift into the video frame of the specific anchor character. Specifically, the media server determines the position of the target anchor character in the video frame based on the mapping relationship between the position information of the face image and the identity identifier established previously. Then, various well-known computer graphics techniques are used to accurately superimpose the virtual gift animated image on the face image of the target anchor character.

[0104] In some embodiments, the synthesis process may involve operations such as layer blending, transparency adjustment, and position correction to ensure that the virtual gift animation image is seamlessly integrated with the live video frame. For example, if the virtual gift is a flower, the animation image will show the flower rotating around the anchor character or slowly falling onto the anchor character. If the gift is a banner, it may unfold on one side of the anchor character or float over the anchor character.

[0105] After the synthesis is completed, the media server distributes the processed live stream to the terminal devices of the users in the live broadcast room. When the audience users watch the live broadcast, they will see the virtual gift animation image added to the live video in real time, which enhances the interactivity and viewing experience of the live broadcast.

[0106] This embodiment not only realizes the real-time recognition and response of multiple identities in the live stream, but also further realizes the deep application of the mapping relationship data between the face image and its identity identifier, thereby bringing significant technical advantages, including but not limited to:

[0107] First, by mapping the location information of the face image with the identity tag, it is possible to quickly and accurately identify the specific anchor person that the viewer clicked in the video frame. This fast matching capability greatly enhances the interactivity of the live broadcast, allowing the viewer's operation instructions to be responded to immediately, improving the user experience.

[0108] Secondly, by responding to the audience's touch operation instructions and popping up the virtual gift selection interface, it provides users with an intuitive way to express their support and love for the anchor. Users can easily select and give virtual gifts. This interaction not only increases the fun of the live broadcast, but also creates an additional source of income for the platform. In this way, the audience's participation and loyalty are improved, and the connection between the anchor and the audience is strengthened.

[0109] Furthermore, by synthesizing virtual gift animation images into live video, a new visual experience is provided to the audience. This synthesis technology makes the virtual gifts look like part of the live scene, enhancing the realism and immersion of gift giving. The audience can see their gifts appear in the host's video in real time. This instant visual feedback not only makes the audience feel more involved, but also adds more entertainment elements to the live event.

[0110] Finally, combined with other embodiments, this embodiment has a high degree of adaptability and flexibility. Whether the number of anchors changes, the audience's interactive needs are diversified, or the live broadcast environment is complex and changeable, it can work stably to ensure the continuity and accuracy of identity recognition and response operations in the live broadcast stream. This stability and technical robustness provide strong technical support for the live broadcast platform, enabling it to stand out in the fiercely competitive market.

[0111] On the basis of any embodiment of the method of the present application, according to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image, responding to the operation instruction acting on the identity of the target anchor character includes:

[0112] Step S3331, responding to the beautification operation instruction, determining the identity identifier and beautification parameters of the target anchor character identity acted upon by the instruction;

[0113] During the live broadcast process, the anchor user often needs to beautify the anchor character in the live video to enhance the visual experience of the live broadcast. This embodiment allows the anchor user to achieve this goal through beauty operation instructions. Specifically, the anchor user can select and set beautification parameters through the beauty interface provided in the live broadcast room. These parameters may include beauty effects, such as skin smoothing, whitening, face slimming, etc., and may also include jewelry decorations, such as wreaths, crowns, etc. Each anchor character can apply unified beautification parameters, or can apply personalized beautification parameters according to personal style or audience preferences.

[0114] When the anchor user selects specific beautification parameters in the beauty interface, these parameters will be recorded and associated with the corresponding anchor character identity. Once the anchor user submits the information set and triggers the beauty operation instruction, the terminal device or media server responsible for implementing the beautification process can determine the identity and beautification parameters of the target anchor character identity affected by the instruction based on these preset or real-time parameters.

[0115] For example, if a host user wants to apply a beauty effect to a female host in the live broadcast room, he or she can select a specific beauty level and decorative accessories in the beauty interface. The background associates these parameters with the identity of the female host and applies these beautification parameters in real time in the live broadcast stream. When the host user issues a beauty operation instruction, the female host's face image in the video frame in the live broadcast stream will be identified and beautified according to the previously set beautification parameters.

[0116] In one embodiment, the anchor user can select different beautification parameters through sliders, drop-down menus or other controls on the beauty interface of the graphical user interface. These parameters will be saved and called when needed to perform real-time processing on the face images in the live video. The device responsible for implementing the image beautification processing will accurately apply the beautification effect to the correct anchor character based on the location information and identity of each face image, ensuring the naturalness and coordination of the live broadcast picture.

[0117] In another embodiment, considering that each anchor person may be an anchor user of their personal live broadcast room, and they usually have their own preferred beautification parameters, which are stored in their personal accounts together with their identity identifiers. In a multi-person live broadcast activity, the anchor user can choose whether to enable the personalized beautification parameters of each anchor person. When this option is enabled, after automatically identifying the anchor person identity identifier corresponding to each face image, the corresponding beautification parameters can be automatically retrieved from their personal accounts.

[0118] For example, in a multi-person live event, suppose there are three anchors, each of whom has his or her own personal account and presets his or her favorite beauty effects and accessories in the account. Before the live broadcast begins, the anchor user of the event can decide whether to enable personalized beautification settings for each participant. Once enabled, after identifying the identity of each face image, the corresponding beautification parameters are automatically extracted from their personal accounts and applied to the live video. In this way, each anchor character will show the beautification effect they are accustomed to in the live broadcast, enhancing the personalized experience of the live broadcast.

[0119] Step S3332: query the mapping relationship data according to the identity identifier, determine the position of the face image of the target anchor character in the video frame, apply the beautification parameters to the face image, and obtain a beautified image;

[0120] When the active anchor user specifies the identity identifier and beautification parameters of the target anchor character identity for the beautification operation instruction, the mapping relationship data can be queried according to the identity identifier to determine the facial image position of the target anchor character identity in the video frame.

[0121] The construction of mapping relationship data is based on the results of face detection and identity matching, where the position of each face image is determined in the video frame through a selection box and associated with the corresponding identity. When the anchor user submits a beautification operation instruction, these mapping relationship data are queried through the identity to quickly locate the exact position of the target anchor character's face image in the video frame.

[0122] After the location is determined, the previously selected beautification parameters are applied to the face image. The process of applying the beautification parameters may include a variety of image processing techniques, such as color correction, filter application, geometric transformation, etc. For example, if the beautification parameters include beauty effects, the target face image may be subjected to skin tone adjustment, skin smoothing, or contour modification. If ornament decoration is included, virtual ornaments, such as a wreath, a crown, etc., may be added to the face image.

[0123] In specific implementation, the application of beautification parameters can be achieved through different algorithms and technologies. For example, a deep learning-based image generation network can be used to generate high-quality beautified images, or traditional image processing techniques can be used to achieve real-time beautification effects. The choice of these technologies depends on the performance requirements of the live broadcast system and the available computing resources. After the beautification parameters are applied, a beautified image corresponding to the face image will be generated.

[0124] Step S3333: Replace the facial image of the target anchor character in the corresponding video frame with the beautified image, so that the video frame displays the corresponding beautified image for the target anchor character in the graphical user interface.

[0125] After obtaining the beautified image corresponding to the facial image of the target anchor, the facial area of ​​the target anchor in the video frame can be accurately identified according to the facial image position information previously determined in the mapping relationship data. Then, the beautified image obtained after applying the beautification parameters is synthesized with the original video frame to replace the original facial image. The specific synthesis process can be achieved through computer graphics technologies such as layer blending, transparency adjustment and position correction to ensure that the beautified image is seamlessly integrated with other parts of the video frame.

[0126] For example, if the anchor character selects skin smoothing and whitening beauty effects during the live broadcast, these effects are applied to the detected face image and a new beautified image is generated. Then, based on the position information of the face image, the corresponding face area is found in the video frame and the original image is replaced with the beautified image. If the anchor character also selects virtual accessories, such as a crown or a wreath, these accessories will also be added to the beautified image and displayed at the corresponding position in the video frame.

[0127] In implementation, this step can adopt a variety of technical solutions. For example, a deep learning-based image generation network can be used to generate high-quality beautified images, or traditional image processing techniques such as filters and color correction can be used to achieve real-time beautification effects. The choice of these technologies depends on the performance requirements of the live broadcast system and the available computing resources.

[0128] After the replacement is completed, the target anchor character in the video frame will appear in the viewer's interface with a beautified image, which improves the visual appeal of the live broadcast and the audience's viewing experience. This real-time beautification process not only enhances the interactivity and viewing value of the live broadcast content, but also provides the anchor with an opportunity to show a personalized image, thereby increasing the appeal of the live broadcast and the audience's participation. In this way, the live broadcast platform can provide a richer and more personalized viewing experience to meet the needs and preferences of different users.

[0129] This embodiment has achieved unique and significant technical advantages over other embodiments of the present application, especially in terms of improving the interactivity and personalized experience of live broadcasts. By allowing anchor users to select and apply personalized beautification parameters for each anchor character participating in a multi-person live broadcast, this embodiment not only meets the anchor's personalized needs for image display, but also provides viewers with a more diverse and rich visual experience. This personalized beautification process allows each anchor to appear in the live broadcast with the image they are most satisfied with, enhancing the attractiveness of the live broadcast and the audience's interest in watching.

[0130] In addition, this embodiment retrieves and applies preset beautification parameters from the personal account of the anchor in an automated manner, greatly simplifying the beautification operation during the live broadcast. The anchor user does not need to manually adjust the beautification effect of each person during the live broadcast, so he can focus more on the creation and interaction of the live broadcast content, improving the efficiency and fluency of the live broadcast. At the same time, this automated application also reduces the errors and inconsistencies caused by manual operations, ensuring the professionalism and consistency of the live broadcast screen.

[0131] More importantly, this embodiment achieves seamless editing and enhancement of the live stream by replacing the original face image in the video frame with the beautified image in real time. This real-time image processing technology not only improves the viewing experience of the live broadcast, but also creates more business opportunities for the live broadcast platform. For example, the platform can attract more users by providing personalized beautification services, or increase revenue by selling virtual accessories and special effects. At the same time, this technology also provides more creative space for anchors. They can customize their live broadcast image according to their own style and the preferences of the audience, thereby enhancing the interactivity and participation of the live broadcast.

[0132] In summary, this embodiment, through its unique technical solution, not only improves the visual quality and personalization level of live broadcast, but also creates more value and opportunities for live broadcast platforms and anchors, thereby achieving significant technical advantages in the live broadcast field.

[0133] On the basis of any embodiment of the method of the present application, solving the optimal match between the face image and the face template based on the distance matrix so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity of the anchor character to which the face template uniquely matches each face image belongs, including:

[0134] Step S3210, subtracting the minimum value of the elements in the corresponding row from each element in each row of the distance matrix to update the distance matrix, and then subtracting the minimum value of the elements in the corresponding column from each element in each column of the distance matrix to update the distance matrix;

[0135] After ensuring that the number of rows and columns of the distance matrix is ​​equal, the matrix can be normalized to provide suitable input for the subsequent matching algorithm. Specifically, each element in the distance matrix needs to be adjusted to reduce the absolute value range of the values ​​in the matrix, thereby simplifying the subsequent matching calculations.

[0136] Specifically, each row of the distance matrix is ​​first processed by subtracting the minimum element value of each row from the elements of each row. This operation does not change the relative differences between the elements within the row, but ensures that the minimum value of each row is zero. The advantage of this is that when looking for lines covering zero values ​​in subsequent steps, it is easier to identify which rows contain zero values. Similarly, the same operation is performed on each column of the distance matrix, subtracting the minimum element value of each column from the elements of each column to ensure that the minimum value of each column is zero.

[0137] In this way, the distance matrix is ​​converted into a new matrix in which the minimum value of each row and column is zero. This normalization process helps to simplify the subsequent matching algorithm because the zero value can be used to determine the optimal match. The zero value indicates that the cost between the face image and the face template is zero, that is, they are optimally matched. For example, suppose there is a 6x6 distance matrix containing the distance values ​​between the real-time face set and the face template in the live character library. After performing the minimum value subtraction operation of the rows and columns, each row and each column in the matrix will have at least one zero value. These zero values ​​will serve as the basis for finding matches in subsequent steps. This normalization process, for the optimal matching algorithm implemented in this embodiment, not only reduces the complexity of numerical calculations, but also provides a clear reference point for the algorithm.

[0138] Step S3220, searching for the minimum number of horizontal and vertical lines that can cover all zero-value elements in the distance matrix. When the total number of lines is the same as the total number of face images, stop searching to determine the distance matrix at this time as the optimal matching matrix. Otherwise, determine the minimum value among the element values ​​not covered by the lines, add the minimum value to the element value corresponding to the intersection of any two lines, and subtract the minimum value from the element value not covered by the lines, and then iterate this step;

[0139] After determining that the number of rows and columns of the distance matrix is ​​equal and normalizing it, the next step is to find a method to determine the optimal match between the face image and the face template. In this embodiment, this optimal match refers to the minimum match. To this end, it is necessary to find the minimum number of horizontal and vertical lines that can cover all zero-valued elements in the distance matrix. These lines represent possible matching paths, and their goal is to minimize the total cost of matching while ensuring that each face image matches a face template, that is, to minimize the sum of data distances.

[0140] Specifically, the normalized distance matrix is ​​first checked to find all zero-valued elements and try to cover these zero values ​​with the least number of lines. Each line, whether horizontal or vertical, represents a potential match between a face image and a face template. When the number of lines found is equal to the total number of face images, it means that a matching face template has been found for each face image, and the distance matrix is ​​considered to be the optimal matching matrix.

[0141] However, if it is found that the number of lines used has not reached the total number of face images in the process of trying to cover all zero-valued elements, this indicates that a matching face template has not been found for each face image. In this case, the distance matrix needs to be further adjusted to promote more matches. The specific method is to find the minimum value among the elements not covered by the lines, add this minimum value to the element corresponding to the intersection of any two lines, and subtract this minimum value from the elements not covered by the lines, and then continue to iterate this step. Such adjustments help find more matching paths in the next iteration.

[0142] In this way, the distance matrix can be iteratively adjusted until the minimum number of lines that can cover all zero-value elements is found, and the number of lines is equal to the number of face images, thereby determining the optimal match between each face image and the face template. The distance matrix at this time can be used as the subsequent optimal matching matrix. This process not only ensures the accuracy of the match, but also improves the matching efficiency because it avoids unnecessary complex calculations and directly optimizes for the goal of finding the optimal solution.

[0143] Step S3230: Based on the optimal matching matrix, find and use augmenting paths to continuously update the matching status between the face image and the face template until no new augmenting paths can be found and the optimal matching relationship between each face image and its corresponding face template is fixed.

[0144] The optimal matching matrix is ​​only the basis for determining the optimal matching, not the optimal matching relationship itself. Therefore, after constructing the optimal matching matrix, continue to use this matrix to find the optimal matching relationship between the face image and the face template. To this end, you can find the augmenting path in this matrix and iteratively update the matching state in the optimal matching matrix until no new augmenting path can be found, thereby determining the optimal matching relationship.

[0145] An augmenting path is a path from a face image to a face template in the optimal matching matrix, where each node on the path alternates between unmatched and matched states. By finding such a path, the matching state can be updated so that more face images can be matched with the corresponding face templates. Specifically, when an augmenting path is found in the optimal matching matrix, the matching state of the face image and face template on the path can be changed, thereby increasing the number of matches.

[0146] First, check the optimal matching matrix to find zero-value elements that can form an augmenting path. If a zero value is found in the matrix, and the row and column where the zero value is located have not been matched, then an augmenting path is found. Along this path, the face image can be matched with the face template, and the corresponding rows and columns are marked as matched. Then, continue to find new augmenting paths in the updated matrix, and repeat this process until no new augmenting paths can be found.

[0147] By continuously searching for and utilizing augmenting paths, the matching status can be gradually updated until all face images in the matrix find matching face templates, thereby fixing a stable matching relationship, that is, each face image establishes a unique matching relationship with its corresponding face template. This is the optimal matching relationship pursued by this embodiment.

[0148] The unique and significant technical advantage that this embodiment brings to this application is that it can effectively handle the situation where the number of face images in the live stream is inconsistent with the preset face template, ensuring the accuracy and efficiency of the matching process. By normalizing the distance matrix, it not only simplifies the subsequent matching calculations, but also provides a clear reference point for the algorithm, so that the matching algorithm can quickly identify the optimal match. This method is particularly suitable for live broadcast environments, where the number of face images may change for various reasons, and this embodiment can dynamically adjust the distance matrix to adapt to these changes.

[0149] In addition, this embodiment iteratively updates the matching state by finding an augmented path. This method can gradually optimize the matching results until the global optimal solution is reached. This iterative optimization process not only improves the accuracy of the matching, but also enhances the adaptability, enabling it to run stably in a complex live broadcast environment. In this way, this embodiment can ensure that each face image can find the face template that best matches it, thereby improving the accuracy of identity recognition in the live stream.

[0150] More importantly, this embodiment can also reduce the consumption of computing resources. Through normalized processing and iterative optimization, unnecessary complex calculations are avoided, making the matching process more efficient. This is especially important for live broadcast platforms that need to process large amounts of data in real time, because it can ensure high performance while also providing a smooth user experience.

[0151] It can be seen that this embodiment, through its innovative technical solution, not only improves the accuracy and efficiency of multi-person identity recognition in live streaming, but also enhances the adaptability and stability of the system, providing a powerful technical support for the live streaming platform. These technical advantages make this application significantly competitive in the live streaming field and provide a solid foundation for improving the interactivity and user experience of live streaming.

[0152] On the basis of any embodiment of the method of the present application, based on the optimal matching matrix, finding and using augmenting paths to continuously update the matching status between the face image and the face template until no new augmenting paths can be found and the optimal matching relationship between each face image and its corresponding face template is fixed, including:

[0153] Step S3231: Create a matching relationship set and initialize the matching relationship set to an empty set;

[0154] In order to record and manage the matching status between each face image and its corresponding face template, a matching relationship set can be created and initialized to be empty in order to subsequently store the matching results obtained by the optimal matching matrix.

[0155] Step S3232: traverse each face image in the optimal matching matrix, find an augmenting path corresponding to one of the face templates for each face image, perform an inversion operation on the augmenting path, and add it to the matching relationship set to update the matching relationship;

[0156] After determining the optimal matching matrix, we can traverse each face image in the matrix and find the corresponding face template for each face image to form the optimal matching relationship. The purpose of detailed analysis of the optimal matching matrix is ​​to find a path that can establish a unique matching relationship between the face image and the face template while ensuring the lowest total matching cost.

[0157] Specifically, for each face image in the matrix, an augmented path is found, that is, a path starting from an unmatched face image, through a series of alternating matched and unmatched states, and finally reaching an unmatched face template. The existence of this path indicates that the number of matches can be increased by changing the current matching state, thereby optimizing the overall matching result.

[0158] After finding such an augmenting path, a key operation is performed: updating the matching relationship. This operation involves reversing the state on the augmenting path, that is, changing the originally matched state on the path to unmatched, and the unmatched state to matched. This state reversal helps to discover new matching opportunities because it may reveal potential matching relationships that were not previously considered.

[0159] After completing this operation, the matching relationship set is updated, and all determined matching relationships can be tracked and managed by adding the results of the augmenting path to the matching relationship set. This process will be repeated until all face images in the optimal matching matrix are traversed and no new augmenting paths can be found.

[0160] Step S3233: The matching relationship set obtained after the traversal is completed is used as the representation of the optimal matching relationship.

[0161] It is not difficult to understand that the process of building the matching relationship set is as follows: every time an augmenting path is found and the matching status is updated, the result of this update is added to the set. As the traversal proceeds, the set gradually becomes richer until all possible augmenting paths have been explored. Finally, the set contains the matching relationships between all face images and the corresponding face templates, and these relationships together constitute a complete representation of the optimal match.

[0162] For example, suppose there are four face images and four face templates in a live broadcast. By traversing and finding augmenting paths, face image 1 may be matched with template 1 first, then face image 2 with template 3, and so on. After each match, these matching relationships are added to the set. When all augmenting paths have been explored and no new augmenting paths can be found, the update stops, and the matching relationship set at this time represents the final optimal matching result.

[0163] This collection not only records the matching results, but also provides a basis for subsequent operations. It can be used as a source table for data representing the mapping relationship between the location information of the face image and the identity identifier of the face template. For example, during the live broadcast, when you need to apply a specific effect to the face image or respond to a specific operation instruction, you can directly query this collection to quickly find the face template corresponding to each face image, and then perform the corresponding operation. This method improves the efficiency of matching and ensures the real-time and interactivity of the live broadcast process.

[0164] The algorithm flow described in this embodiment brings significant technical advantages to this application, especially in ensuring the efficiency of multi-person identity recognition in live streaming. By creating and utilizing a set of matching relationships, the process can systematically record and manage the matching status of each face image and its corresponding face template, thereby achieving fast and accurate matching result updates. The core of this method lies in its ability to iteratively find augmented paths, which allows the best possible match to be found at each step, and continuously optimizes the matching relationship through inversion operations until the global optimal solution is reached.

[0165] This update mechanism based on augmented paths not only improves the flexibility of the matching process, but also greatly improves the matching efficiency. In a live broadcast environment, real-time performance is crucial, and the method described in this embodiment can ensure the accuracy and stability of the matching results while maintaining real-time performance. By traversing and searching for augmented paths in the optimal matching matrix, it is possible to quickly respond to changes in the live stream, such as newly appeared faces or face templates, thereby dynamically adjusting the matching relationship set to ensure that each face image can be correctly matched with its corresponding face template.

[0166] In addition, the efficiency of this process is also reflected in its optimal use of computing resources. By avoiding unnecessary complex calculations, a large amount of data can be processed in a short time, which is particularly important for live broadcast platforms that need to process high frame rate video streams. This method reduces the computing burden, allowing more resources to be invested in improving the quality of live broadcasts and enhancing the user interaction experience.

[0167] In summary, this embodiment not only improves the accuracy and real-time performance of multi-person identity recognition in live streaming, but also optimizes the use of computing resources, enabling the live streaming platform to provide a more fluent and interactive live streaming experience. These advantages make this application significantly competitive in the field of live streaming technology and provide a solid technical foundation for the future development of live streaming platforms.

[0168] See also Figure 3 According to one aspect of the present application, a device for identifying multiple people in a live stream is provided, comprising a similarity operation module 3100, a face matching module 3200, and a mapping response module 3300, wherein the similarity operation module 3100 is configured to determine the data distance between each face image in the live stream of a live broadcast room and a plurality of face templates corresponding to a plurality of anchor characters in a preset live broadcast character library corresponding to the live broadcast room, and obtain a distance matrix; the face matching module 3200 is configured to solve the optimal match between the face image and the face template based on the distance matrix, so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity identifier of the anchor character identity to which the face template uniquely matches each face image belongs; the mapping response module 3300 is configured to respond to the operation instruction acting on the target anchor character identity according to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image.

[0169] Based on any embodiment of the device of the present application, the similarity operation module 3100 includes: a tracking and detection module, which is configured to extract video frames from the live stream transmitted in real time in the live broadcast room, and perform face tracking and detection on the video frames to obtain multiple face images therein to form a real-time face set; a distance calculation module, which is configured to calculate the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and express the data distance relationship between each face image and each face template as a distance matrix; a row and column regularization module, which is configured to detect whether the number of rows and columns of the distance matrix is ​​the same. When the two are different, the missing rows or columns are supplemented, and each element therein is assigned a zero value to update the distance matrix.

[0170] Based on any embodiment of the device of the present application, the similarity operation module 3100 includes: a tracking and detection module, which is configured to extract video frames from the live stream transmitted in real time in the live broadcast room, and perform face tracking and detection on the video frames to obtain multiple face images therein to form a real-time face set; a distance calculation module, which is configured to calculate the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and express the data distance relationship between each face image and each face template as a distance matrix; a row and column alignment module, which is configured to calculate the larger number of rows and columns of the distance matrix, and for each row / column with a larger number, calculate the distance mean corresponding to the element value of each column / row, and delete several rows / columns whose distance mean indicates the lowest similarity, so that the number of rows and columns of the distance matrix is ​​consistent.

[0171] Based on any embodiment of the device of the present application, the mapping response module 3300 includes: a gift response module, which is configured to respond to a touch operation instruction acting on the target anchor character identity in the display image of the video frame in the graphical user interface, and determine the identity of the target anchor character identity according to the mapping relationship data between the position information of each facial image in the video frame and the identity of the facial template matching the facial image; a gift determination module, which is configured to pop up a virtual gift selection interface to obtain gift configuration information, which includes the virtual gift determined by the user to be given to the target anchor character identity; a gift synthesis module, which is configured to respond to the gift operation instruction submitted by the user, obtain the animated image of the virtual gift in the gift configuration information, and synthesize the animated image into the facial image of the target anchor character identity corresponding to the identity identifier in each video frame of the live stream.

[0172] Based on any embodiment of the device of the present application, the mapping response module 3300 includes: a beautification response module, which is configured to respond to a beautification operation instruction and determine the identity identifier and beautification parameters of the target anchor character identity acted upon by the instruction; a beautification application module, which is configured to query mapping relationship data according to the identity identifier, determine the position of the facial image of the target anchor character identity in the video frame, apply the beautification parameters to the facial image, and obtain a beautified image; a beautification synthesis module, which is configured to replace the facial image of the target anchor character identity in the corresponding video frame with the beautified image, so that the video frame displays the corresponding beautified image for the target anchor character identity in the graphical user interface.

[0173] Based on any embodiment of the device of the present application, the face matching module 3200 includes: an initial processing module, configured to subtract the minimum value of the elements in the corresponding row from each element in each row of the distance matrix to update the distance matrix, and then subtract the minimum value of the elements in the corresponding column from each element in each column of the distance matrix to update the distance matrix; a flow calculation module, configured to find the minimum number of horizontal lines and vertical lines that can cover all zero-value elements in the distance matrix, and when the total number of lines is the same as the total number of face images, stop searching to determine the distance matrix at this time as the optimal matching matrix, otherwise, determine the minimum value of the element values ​​not covered by the lines, add the minimum value to the element values ​​corresponding to the intersection of any two lines, and subtract the minimum value from the element values ​​not covered by the lines and iterate this step; a matching determination module, configured to find and use the augmentation path based on the optimal matching matrix to continuously update the matching status of the face image and the face template until no new augmentation path can be found and the optimal matching relationship between each face image and its corresponding face template is fixed.

[0174] Based on any embodiment of the device of the present application, the matching determination module includes: a set creation module, configured to create a matching relationship set and initialize the matching relationship set to an empty set; an augmentation update module, configured to traverse each face image based on the optimal matching matrix, and find an augmentation path corresponding to one of the face templates for each face image, and add the augmented path to the matching relationship set after reversing the augmented path to update the matching relationship; a relationship representation module, configured to use the matching relationship set obtained after completing the traversal as a representation of the optimal matching relationship.

[0175] Another embodiment of the present application also provides a device for identifying multiple people in a live stream. Figure 4As shown, a schematic diagram of the internal structure of a device for identifying multiple people in a live stream. The device for identifying multiple people in a live stream includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable non-volatile storage medium of the device for identifying multiple people in a live stream stores an operating system, a database, and computer-readable instructions. The database may store an information sequence. When the computer-readable instructions are executed by the processor, the processor may implement a method for identifying multiple people in a live stream.

[0176] The processor of the multi-person identity recognition device in the live stream is used to provide computing and control capabilities to support the operation of the multi-person identity recognition device in the live stream. The memory of the multi-person identity recognition device in the live stream may store computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor may execute the multi-person identity recognition method in the live stream of the present application. The network interface of the multi-person identity recognition device in the live stream is used to connect and communicate with the terminal.

[0177] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the multi-person identity recognition device in a live stream to which the scheme of the present application is applied. The specific multi-person identity recognition device in a live stream may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0178] In this embodiment, the processor is used to execute Figure 3 The memory stores the program code and various data required to execute the above modules or submodules. The network interface is used to realize data transmission between user terminals or servers. The non-volatile readable storage medium in this embodiment stores the program code and data required to execute all modules in the multi-person identity recognition device in the live stream of this application, and the server can call the program code and data of the server to execute the functions of all modules.

[0179] The present application also provides a non-volatile readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the method for identifying multiple people in a live stream of any embodiment of the present application.

[0180] The present application also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any embodiment of the present application when executed by one or more processors.

[0181] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0182] In summary, this application can accurately handle common problems such as side faces, lowered heads, occlusions, exaggerated expressions, and motion blur in live broadcast scenes, ensuring accurate identity recognition in multi-person live broadcast environments. By constructing and optimizing the distance matrix, this application not only improves the accuracy of face recognition, but also achieves global optimal matching, avoiding the mismatch and duplication of host character identities that may occur when traditional methods deal with multi-person live broadcast scenes. In addition, this application establishes an identity mapping relationship between each face image and the corresponding face template, so that the system can respond to user operation instructions for specific characters, such as giving virtual gifts and applying beauty effects, thereby greatly enhancing the interactivity and fun of live broadcasts. These advantages not only improve the user experience, but also provide more personalized and interactive tools for live content creators, further enriching the live broadcast gameplay and promoting the innovation and development of live broadcast technology.

Claims

1. A method for identifying multiple people in a live stream, characterized in that: include: Determine the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room, and obtain a distance matrix; Solve the optimal match between the face image and the face template based on the distance matrix, so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity of the anchor character to which the face template uniquely matches each face image belongs; According to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image, the operation instruction acting on the identity of the target anchor character is responded to.

2. The method for identifying multiple people in a live stream according to claim 1, characterized in that: Determine the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room, and obtain a distance matrix, including: Extracting video frames from the live stream transmitted in real time in the live broadcast room, performing face tracking detection on the video frames to obtain multiple face images therein to form a real-time face set; Calculate the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and express the data distance relationship between each face image and each face template as a distance matrix; Check whether the number of rows and columns of the distance matrix is ​​the same. If they are different, add the missing rows or columns and assign zero values ​​to each element therein to update the distance matrix.

3. The method for identifying multiple people in a live stream according to claim 1, characterized in that: Determine the data distance between each face image in the live stream of the live broadcast room and multiple face templates corresponding to multiple anchor characters in the preset live broadcast character library corresponding to the live broadcast room, and obtain a distance matrix, including: Extracting video frames from the live stream transmitted in real time in the live broadcast room, performing face tracking detection on the video frames to obtain multiple face images therein to form a real-time face set; Calculate the data distance between the semantic features of each face image in the real-time face set and each face template in the live character library, and express the data distance relationship between each face image and each face template as a distance matrix; Calculate the larger number of rows and columns of the distance matrix, and for each row / column with a larger number, calculate the mean distance corresponding to the element values ​​of each column / row, and delete the rows / columns whose distance means represent the lowest similarity, so that the number of rows and columns of the distance matrix is ​​consistent.

4. The method for identifying multiple people in a live stream according to claim 1, characterized in that: According to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image, respond to the operation instruction acting on the identity of the target anchor person, including: In response to a touch operation instruction acting on a target anchor character identity in a display image of a video frame in a graphical user interface, an identity identifier of the target anchor character identity is determined according to mapping relationship data between position information of each face image in the video frame and an identity identifier of a face template matching the face image; A virtual gift selection interface is popped up to obtain gift configuration information, which includes the virtual gift determined by the user to be given to the target anchor character; In response to a gift giving operation instruction submitted by a user, an animated image of a virtual gift in the gift giving configuration information is obtained, and the animated image is synthesized into a face image of a target anchor character corresponding to the identity identifier in each video frame of the live stream.

5. The method for identifying multiple people in a live stream according to claim 1, characterized in that: According to the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image, respond to the operation instruction acting on the identity of the target anchor person, including: Responding to a beautification operation instruction, determining the identity identifier and beautification parameters of the target anchor character identity acted upon by the instruction; Querying the mapping relationship data according to the identity identifier, determining the position of the face image of the target anchor character in the video frame, and applying the beautification parameters to the face image to obtain a beautified image; The facial image of the target anchor character identity in the corresponding video frame is replaced with the beautified image, so that the video frame displays the corresponding beautified image for the target anchor character identity in the graphical user interface.

6. The method for identifying multiple people in a live stream according to any one of claims 1 to 5, characterized in that: Solving the optimal match between the face image and the face template based on the distance matrix so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity of the anchor character to which the face template uniquely matches each face image belongs, including: Subtracting the minimum value of the elements in the corresponding row from each element in each row of the distance matrix to update the distance matrix, and then subtracting the minimum value of the elements in the corresponding column from each element in each column of the distance matrix to update the distance matrix; Find the minimum number of horizontal and vertical lines that can cover all zero-value elements in the distance matrix. When the total number of lines is the same as the total number of face images, stop searching to determine the distance matrix at this time as the optimal matching matrix. Otherwise, determine the minimum value of the element values ​​not covered by the lines, add the minimum value to the element value corresponding to the intersection of any two lines, and subtract the minimum value from the element value not covered by the lines, and then iterate this step; Based on the optimal matching matrix, an augmenting path is found and used to continuously update the matching status between the face image and the face template until no new augmenting path can be found and the optimal matching relationship between each face image and its corresponding face template is fixed.

7. The method for identifying multiple people in a live stream according to claim 6, characterized in that: Based on the optimal matching matrix, finding and using augmenting paths to continuously update the matching status between the face image and the face template until no new augmenting paths can be found and the optimal matching relationship between each face image and its corresponding face template is fixed, including: Create a matching relationship set and initialize the matching relationship set to an empty set; Based on traversing each face image in the optimal matching matrix, finding an augmenting path corresponding to one of the face templates for each face image, and adding the augmenting path to the matching relationship set after performing an inversion operation to update the matching relationship; The matching relationship set obtained after the traversal is completed is used as the representation of the optimal matching relationship.

8. A device for identifying multiple people in a live stream, characterized in that: include: A similarity operation module is configured to determine the data distance between each face image in the live stream of the live broadcast room and a plurality of face templates corresponding to a plurality of anchor characters in a preset live broadcast character library corresponding to the live broadcast room, and obtain a distance matrix; A face matching module is configured to solve the optimal match between the face image and the face template based on the distance matrix, so that the sum of the data distances between the face image and the face template is the optimal solution, so as to determine the identity of the anchor character to which the face template uniquely matches each face image belongs; The mapping response module is configured to respond to the operation instructions acting on the identity of the target anchor character based on the mapping relationship data between the position information of each face image in the video frame and the identity identifier of the face template matching the face image.

9. A device for identifying multiple people in a live stream, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A non-volatile readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

Citation Information

Cited By

  • Live broadcast interaction method and device

    CN120378642A

  • Character recognition method and device for live broadcast video, live broadcast system, equipment and medium

    CN120783376A

  • Method and apparatus for multi-person identity recognition in live stream, and device and medium

    WO2026157463A1