Data analysis method of user group and related device
Through multi-terminal collaborative data collection and cloud-based intelligent analysis, and using image algorithms to identify audience expressions and movements, the problem of low accuracy in manual observation is solved, and automated, comprehensive and accurate analysis of user group feedback data is achieved.
Patent Information
- Application Number
- CN202511161625.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-19
AI Technical Summary
In the existing technology, the accuracy of obtaining audience feedback data by relying on manual observation is low, and there are problems such as poor real-time performance, limited perspective, and single data analysis dimension, making it difficult to fully capture the real reactions of the user group.
It adopts a multi-terminal collaborative collection and cloud-based intelligent analysis solution, uses image algorithms to recognize audience expressions and movements in real time, obtains key frames in video segments, identifies facial and body key point information, performs location division and user tracking, and determines action information.
It improves the accuracy of data recognition, avoids the influence of human subjective factors, realizes automatic recognition of user group actions and emotions, and enhances the comprehensiveness and accuracy of data analysis.
Smart Images

Figure CN120673347A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis, and more specifically, to a method and related device for analyzing user group data. Background Art
[0002] In large-scale settings like concerts and training sessions, the audience or trainees can be referred to as a user group. These user groups are typically quite large. In these scenarios, the user group's experience is crucial to the effectiveness of the concert or training session.
[0003] Currently, taking concerts as an example, audience feedback on performances is obtained through manual observation. For example, a director in the audience can determine which audience members are in high spirits based on their subjective experience, and then use a camera to capture videos of these audience members in high spirits. However, this method of determining feedback data such as audience emotions is easily affected by factors such as human subjectivity, resulting in low accuracy of the determined audience feedback data. Summary of the Invention
[0004] In view of this, the present application provides a user group data analysis method and related devices to solve the problem of low accuracy in determining feedback data.
[0005] To solve the above technical problems, this application adopts the following technical solutions:
[0006] A method for analyzing user group data, comprising:
[0007] Acquire key frames in a video segment collected by at least one data collection terminal at a current moment; each of the data collection terminals is used to collect video information of users in different areas of a user group;
[0008] Identifying facial information and body key point information corresponding to the same user in the key frame; the facial information at least includes emotional information;
[0009] Based on the facial information of the user, performing a location division operation on the user's location to determine target users corresponding to the same location attribute;
[0010] Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result;
[0011] Based on at least one of facial information and body key point information of the target user in the user tracking result at different moments, action information of the target user in the user tracking result is determined.
[0012] Optionally, identifying facial information and body key point information corresponding to the same user in the key frame includes:
[0013] Identify the user's facial information and body key point information in the key frame; the facial information includes facial emotion information and facial key point information in the face detection frame;
[0014] Based on the relative distance between the face detection frame and the designated key point in the human body key point information, a matching operation is performed on the face information and the human body key point information to obtain the face information and the human body key point information corresponding to the same user.
[0015] Optionally, based on the facial information of the user, performing a location division operation on the user's location to determine target users corresponding to the same location attribute includes:
[0016] performing a sorting operation on the users based on the center heights of the face detection frames in the face information to obtain a sorting result;
[0017] Based on the average height of the face detection frame and the average center height of the face detection frame, a position classification operation is performed on the users in the sorting result to determine target users belonging to the same position attribute; the position attribute includes a sorting number.
[0018] Optionally, based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result, including:
[0019] Calculating the difference between the horizontal coordinate of each target user in the key frame at the current moment and the horizontal coordinate of each target user in the key frame at the previous moment;
[0020] Matching the target user in the key frame at the current moment with the target user in the key frame at the previous moment in ascending order of the difference values to obtain a user matching result;
[0021] For a user to be assigned that is not in the user matching result among the target users in the key frame at the current moment, determining position information of the user to be assigned based on the horizontal coordinates of the user to be assigned;
[0022] A user tracking result is determined based on the user matching result and the location information of the user to be assigned.
[0023] Optionally, after determining the action information of the target user in the user tracking result, the method further includes:
[0024] The seat number, play timestamp, emotion information and action information of the target user are stored in the target area.
[0025] Optionally, after storing the seat number, playback timestamp, emotion information, and action information of the target user in the target area, the method further includes:
[0026] Pulling the latest data stored in the target area;
[0027] If the playback timestamp in the latest data is less than or equal to a preset time after the pulling time, at the preset time, data of a specified time period stored in the target area is extracted; the specified time period is the time period between the pulling time and the preset time.
[0028] Optionally, after extracting the data of the specified time period stored in the target area, the method further includes:
[0029] When the number of users in the latest data is greater than a preset number, determining current emotion distribution data of the user group in the latest data;
[0030] Based on the current emotion distribution data, performing an emotion information filling operation on users in the user group for whom emotion information does not exist;
[0031] When the number of users in the latest data is not greater than a preset number, default emotion distribution data is used to perform an emotion information filling operation on users in the user group for whom emotion information does not exist.
[0032] A data analysis device for a user group, comprising:
[0033] An acquisition module, configured to acquire key frames from a video segment currently acquired by at least one data acquisition terminal; each of the data acquisition terminals being configured to acquire video information of users located in different areas of a user group;
[0034] A recognition module, configured to identify facial information and body key point information corresponding to the same user in the key frame; the facial information at least includes emotional information;
[0035] a segmentation module, configured to perform a location segmentation operation on the user's location based on the user's facial information, so as to determine target users corresponding to the same location attribute;
[0036] a matching module, configured to match the target user in the key frame at the current moment with the target user in the key frame at the previous moment based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, to obtain a user tracking result;
[0037] The determination module is configured to determine the action information of the target user in the user tracking result based on at least one of the facial information and the body key point information of the target user at different moments in the user tracking result.
[0038] An electronic device comprising at least one processor and a memory connected to the processor, wherein:
[0039] The memory is used to store computer programs;
[0040] The processor is used to execute the computer program so that the electronic device can implement the above-mentioned user group data analysis method.
[0041] A computer storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the above-mentioned user group data analysis method.
[0042] The present application provides a data analysis method and related devices for a user group. In the present application, at least one data acquisition terminal is used to collect video information of users in different areas of the user group, which can cover users in all areas and ensure the comprehensiveness of user analysis. Subsequently, the facial information and body key point information corresponding to the same user in the key frame are identified. Based on the facial information of the user, a position division operation is performed on the location of the user to determine the target user corresponding to the same position attribute. Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result. Based on at least one of the facial information and body key point information of the target user in the user tracking result at different moments, the action information of the target user in the user tracking result is determined. That is, in the present application, the user's actions and emotions can be automatically identified based on the above process, avoiding the influence of human subjective factors on the recognition results and improving the accuracy of data recognition. In addition, during action recognition, the target user's facial information at different times and at least one of the body key point information in the user tracking results are determined. Since the body key points at different times are taken into account, that is, the user's posture at different times is taken into account, the accuracy of user action determination is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0044] Figure 1 A flowchart of a method for analyzing user group data provided in an embodiment of the present application;
[0045] Figure 2 A flowchart of a user determination method provided in an embodiment of the present application;
[0046] Figure 3 A user tracking flowchart provided in an embodiment of the present application;
[0047] Figure 4 A data pulling flow chart provided in an embodiment of the present application;
[0048] Figure 5 A schematic diagram of a scenario of a user group data analysis method provided in an embodiment of the present application;
[0049] Figure 6 A schematic diagram of the structure of a user group data analysis device provided in an embodiment of the present application;
[0050] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] In large-scale scenarios like variety shows, concerts, and training sessions, the audience or trainees can be referred to as a user group. These user groups are typically quite large. In these scenarios, the user group's experience is crucial to the effectiveness of the concert or training session.
[0053] Currently, with the booming market for variety shows and concerts, audience interaction has become a key factor in enhancing program effectiveness and commercial value. Audience feedback analysis primarily relies on manual observation. For example, a director in the audience can identify audience members with high emotions based on their subjective experience. A camera can then capture video of these audience members, which can then be edited to produce the corresponding video. This method of determining audience feedback data, such as emotions, is susceptible to subjective factors, resulting in low accuracy. Furthermore, it suffers from issues such as poor real-time performance, limited viewing angles, and a single-dimensional data analysis (e.g., analyzing only emotions but not actions), making it difficult to fully capture the audience's true reactions.
[0054] Recent advances in artificial intelligence and cloud computing have opened up new possibilities for analyzing audience behavior. Computer vision technology can identify audience expressions and movements in real time, providing data support for program production. However, existing solutions still face numerous challenges. For one thing, the limited field of view of a single camera leads to serious audience occlusion. Furthermore, complex on-site environments (such as dense crowds) significantly reduce recognition accuracy.
[0055] To this end, in this application, in order to accurately identify audience feedback data in situations where the audience is dense and the viewing angle is limited, a multi-terminal collaborative collection and cloud-based intelligent analysis solution is adopted. Image algorithms (face recognition, face key point recognition, face emotion recognition, and body key point recognition) are used to analyze the audience's expressions / movements at the concert in real time, and the results are synchronized to the live broadcast page in real time. When watching the singer's performance, the audience can immerse themselves in the atmosphere of the scene. At the same time, the singer's performance can be scored based on the number and frequency of emotional resonance achieved during the performance, which is more objective and fair than manual scoring.
[0056] Specifically, the present application provides a data analysis method and related devices for a user group. In the present application, at least one data acquisition terminal is used to collect video information of users in different areas of the user group, which can cover users in all areas and ensure the comprehensiveness of user analysis. Subsequently, the facial information and body key point information corresponding to the same user in the key frame are identified. Based on the facial information of the user, a position division operation is performed on the location of the user to determine the target user corresponding to the same position attribute. Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result. Based on at least one of the facial information and body key point information of the target user in the user tracking result at different moments, the action information of the target user in the user tracking result is determined. That is, in the present application, the user's actions and emotions can be automatically identified based on the above process, avoiding the influence of human subjective factors on the recognition results and improving data recognition accuracy. In addition, during action recognition, the target user's facial information at different times and at least one of the body key point information in the user tracking results are determined. Since the body key points at different times are taken into account, that is, the user's posture at different times is taken into account, the accuracy of user action determination is improved.
[0057] In one implementation, referring to Figure 1 , a data analysis method for a user group may include:
[0058] S11. Obtain key frames in video segments collected by at least one data collection terminal at a current moment.
[0059] Each of the data collection terminals is used to collect video information of users in different areas of the user group.
[0060] In specific implementation, taking a concert as an example, the audience area is divided into multiple zones. The audience numbers in each zone are numbered and recorded in order to create a seating map. Terminals are set up on-site and the image acquisition parameters and angles are adjusted. Multiple terminals (such as mobile phones and cameras) are positioned in different areas in front of the auditorium. The combined coverage area of these terminals should cover the entire concert audience area. Subsequently, footage of the audience area can be captured, with the angle ensuring that each row of audience members is parallel to the horizon. The terminals use an app to send the captured video to a live streaming cloud server via the RTMP (Real-Time Messaging Protocol). In addition to RTMP for data transmission, other formats, such as RTC (Real-Time Communication) and m3u8, can also be used, depending on the specific configuration, to ensure consistent streaming latency and stable image quality.
[0061] For each terminal, the live cloud server receives the data transmitted by the terminal in real time, and uses the live stream transfer process to transfer the network live stream uploaded by the terminal into a local m3u8 format video file list, and records the timestamp Stream_start_time when the transfer starts. When the video stream is interrupted and restored due to various reasons, the current process needs to automatically restart the transfer and record the timestamp.
[0062] Use the read analysis process to continuously read the local m3u8 format video file list and calculate the timestamp of each video segment. The method for calculating the timestamp is:
[0063] If the write time of the current m3u8 file list is 8:00:00 and the duration of each video segment is 2 seconds, then the timestamps of the 1st, 2nd, 3rd, and 4th video segments are 8:00:02, 8:00:04, 8:00:06, and 8:00:08 respectively. Then, parse the key frames in the video segments.
[0064] It should be noted that the live stream transfer process and the reading and analysis process in this application are in a multi-copy mode, with one terminal corresponding to one copy.
[0065] S12: Identify facial information and body key point information corresponding to the same user in the key frame.
[0066] The facial information at least includes emotional information, and the emotional information may be happiness, sadness, anger, etc.
[0067] In one implementation, the facial information includes facial emotion information and facial key point information in a face detection frame.
[0068] The model can be used to detect human key point information, facial emotion information within the face detection frame, and facial key point information. After identifying the facial emotion information, facial key point information, and human key point information, it is necessary to establish a correspondence between the facial emotion information, facial key point information, and human key point information, that is, to determine which facial emotion information, facial key point information, and human key point information belong to the same viewer.
[0069] In one implementation, step S12 may include:
[0070] 1) Recognize the user's facial information and body key point information in the key frame.
[0071] The facial information includes facial emotion information and facial key point information in the face detection frame.
[0072] In specific implementation, the face detection model can be used to detect the face area in the key frame, and the face detection frame can be used to frame the face area. The face detection frame is set to a rectangular area centered on the center point of the face, with a height of 4 / 5 of the shoulder width and a width of 3 / 5 of the shoulder width.
[0073] Then, the facial emotion model is used to identify the emotion of the facial area framed by the face detection frame, and the facial key point model is used to identify the key points of the facial area framed by the face detection frame. The key points are called facial key point information, and the facial key point information can be areas such as the nose, eyes, and mouth.
[0074] In addition, the human body key point model can be used to analyze the key frames to obtain human body key point information. The human body key point information can be human body joint key points, such as: head center point, wrist, elbow, shoulder, knee, etc.
[0075] Among them, the model in this application can be used after training a general deep learning model.
[0076] 2) Based on the relative distance between the face detection frame and the designated key points in the human key point information, a matching operation is performed on the face information and the human key point information to obtain the face information and the human key point information corresponding to the same user.
[0077] The designated key point in the human body key point information may be the center point of the head.
[0078] In specific implementation, the center point of each face detection frame is matched with the center point of the head in each human body key point information. If the relative distance between the center point of the face detection frame and the center point of the head is less than a threshold, it is determined that the user corresponding to the face detection frame and the user corresponding to the center point of the head are the same user.
[0079] The threshold can be set to the height of the face detection box bbox_height. For face information that does not match human key point information, its human key point information is empty. For face information that does not match human key point information, its facial information is empty.
[0080] Through this step, the face and body of the same user can be matched, so that the emotions and actions of the same user can be defined.
[0081] S13: Based on the facial information of the user, perform a location division operation on the user's location to determine target users corresponding to the same location attribute.
[0082] In this application, after matching the user's emotions and body key points through the above steps, it is necessary to sort the user into rows. The seats in the auditorium are divided into rows and columns. In this application, sorting the user into rows can determine which row the user belongs to, and then determine which seat the user belongs to within that row.
[0083] Among them, the same position attribute in this application is the same sort number, that is, users located in the same row are screened out, and these users can be called target users.
[0084] In one implementation, referring to Figure 2 , step S13 may include:
[0085] S21. Sorting the users based on the center heights of the face detection frames in the face information to obtain a sorting result.
[0086] In the specific implementation, after obtaining all the face detection frames in a key frame, all the face detection frames are sorted by the center point height of the face detection frame to obtain the sorting result. In the sorting result, the face detection frame with smaller center point height is ranked higher, and the face detection frame with larger center point height is ranked lower.
[0087] In actual scenarios, the center points of the face detection frames of users in the same row have similar heights. Therefore, the center point heights of the face detection frames can be used for subsequent sorting operations.
[0088] S22: Based on the average height of the face detection frame and the average center height of the face detection frame, perform a position classification operation on the users in the sorting result to determine target users belonging to the same position attribute.
[0089] Wherein, the position attribute includes a sequence number.
[0090] In real-world scenarios, the center points of the face detection frames of users in the same row are similar in height. The height difference between the face detection frames of different users in the same row does not exceed the average height of the face detection frames of the users in that row. Therefore, based on this principle, users are sorted.
[0091] In the specific implementation, take the first data in the sorting result, use the center point height of the face detection frame of the first data as the center height mean value center_y_mean of the face frame of the current row, and the height difference of the face detection frame of the first data as the average height bbox_y_mean of the face frame of the current row.
[0092] Then, take the second data from the sorted result. If the difference between the center point height of the face detection frame in the second data and center_y_mean does not exceed bbox_y_mean, it means that the user corresponding to the second data is in the same row as the user corresponding to the first data. Then, take the average of the center point heights of the first user's face detection frame and the second user's face detection frame as the new face frame center height mean value center_y_mean for the current row, and take the average of the height difference between the face detection frame of the first data and the second data as the new face frame average height bbox_y_mean for the current row.
[0093] Then, continue to extract the second data from the sorting results and sort it according to the above process. If the difference between the center point height of the face detection box in the third data and center_y_mean does not exceed bbox_y_mean, the user corresponding to the third data is in the same row as the first two users. If the difference between the center point height of the face detection box in the third data and center_y_mean exceeds bbox_y_mean, the user corresponding to the third data is not in the same row as the first two users. In this case, the user corresponding to the third data is placed in the row after the first two users, and statistics for a new row are started. This process continues until every data point in the sorting results is analyzed, resulting in a list of detection data classified by row.
[0094] In this embodiment, by classifying users by rows, it is possible to subsequently accurately locate the row and seat of the user, thereby achieving accurate positioning of the user's position, and further achieving user tracking at different times and accurate determination of the user's posture.
[0095] S14. Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result.
[0096] In this embodiment, after determining the user's ranking number, the user tracking operation can be implemented based on the key frames of the two previous and subsequent moments, and the same user in the key frames at different moments can be determined, thereby determining the user's action based on the posture of the same user at different moments.
[0097] S15. Determine action information of the target user in the user tracking result based on at least one of facial information and body key point information of the target user at different moments in the user tracking result.
[0098] In this embodiment, the action information may include various action postures such as standing, swaying, and clapping.
[0099] The standing determination process involves calculating the angle between the torso and thigh, and the length ratio between the torso and thigh, based on key point information. The closer the angle is to 180 degrees, the closer the body is to being upright (or facing the camera). When standing, if the torso-thigh length ratio is relatively small, this may be due to sitting facing the camera. The angle and length ratio between the line formed by the left shoulder and left thigh root and the line formed by the left thigh root and left knee joint are calculated. If the angle is greater than 170 degrees and the length ratio is less than 2, the person is considered to be standing on the left side. The same method is used to determine if the person is standing on the right side.
[0100] It should be noted that when judging whether the audience is standing, since only the audience in the front row will block the audience in the back row, the standing judgment needs to be made in the order from the front row to the back row.
[0101] In one implementation, since the legs of back-row spectators are often obscured, a standing mask is drawn for the face and torso detection frames of spectators identified as standing. If the ratio of the intersection of the back-row spectator's torso and the standing mask to the spectator's torso area is greater than 0.05, occlusion is determined, and the spectator is considered standing. For each spectator identified as standing, the standing mask is updated, setting the standing mask for the face and torso detection frames to 1.
[0102] The process for determining whether an applause is occurring is as follows: Applause is a dynamic action. First, the wrist-elbow angle is determined to be between 30 and 60 degrees. If so, the wrist distance is calculated; otherwise, the wrist distance is returned as None. The wrist distances for the last four consecutive moments are then evaluated. If the minimum value of these consecutive non-None wrist distances is less than 1 / 4 of the shoulder length, the applause is considered occurring.
[0103] The process for determining sway is as follows: record the position of the head center (head_center) and the angle between the head and the navel (head_angle). Starting from the second keyframe, calculate the change in the head center position distance between adjacent frames. Calculate the head distance change and head angle change for the last five consecutive moments. If the difference between the maximum and minimum head angles is greater than 15 degrees (angle change), and the head distance change is greater than 0.1 times the shoulder length and less than 1 times the shoulder length (head displacement is present but not significant), sway is determined.
[0104] In addition, for other action postures, there is corresponding judgment logic, which can be used to make action judgments.
[0105] In this embodiment, at least one data acquisition terminal is used to collect video information of users in different areas of the user group, which can cover users in all areas and ensure the comprehensiveness of user analysis. Subsequently, the facial information and body key point information corresponding to the same user in the key frame are identified, and based on the facial information of the user, a position division operation is performed on the location of the user to determine the target user corresponding to the same position attribute. Based on the coordinates of the target user in the key frame of the current moment and the coordinates of the target user in the key frame of the previous moment, a matching operation is performed on the target user in the key frame of the current moment and the target user in the key frame of the previous moment to obtain the user tracking result. Based on at least one of the facial information and body key point information of the target user in the user tracking result at different moments, the action information of the target user in the user tracking result is determined. That is, in this application, the user's actions and emotions can be automatically identified based on the above process, avoiding the influence of human subjective factors on the recognition results and improving the accuracy of data recognition. In addition, during action recognition, the target user's facial information at different times and at least one of the body key point information in the user tracking results are determined. Since the body key points at different times are taken into account, that is, the user's posture at different times is taken into account, the accuracy of user action determination is improved.
[0106] Based on any of the above embodiments, refer to Figure 3 , based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, performing a matching operation on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result, which may include:
[0107] S31 : Calculate the difference between the horizontal coordinate of each target user in the key frame at the current moment and the horizontal coordinate of each target user in the key frame at the previous moment.
[0108] Among them, if the key frame at the current moment is the first key frame, there is no key frame at the previous moment. At this time, according to the stored seat distribution map, the audience in each row can be assigned seat numbers according to the left-right order and their seat positions can be recorded.
[0109] If the current keyframe is not the first keyframe, and there is a keyframe from the previous moment, a seat matching method based on minimum Euclidean distance can be used. By matching the horizontal coordinates of the current row of viewers with the horizontal coordinates of the previous row of viewers, the seat number of each viewer in the current keyframe can be matched. In this application, the seat number of each viewer in the current keyframe is matched because the live broadcast page needs to display expressions / actions at the positions corresponding to the seats, so a seat number needs to be assigned to each expression / action detection result.
[0110] When matching the horizontal coordinates of the current row of spectators with the horizontal coordinates of the row's history using the minimum distance principle, the horizontal coordinates of the row in the current keyframe and the horizontal coordinate list of the row at the previous moment are first converted into two-dimensional NumPy arrays. In real-world scenarios, spectators in the same row may temporarily leave or return, causing the number of spectators in the row in the current keyframe to be the same or different from the number in the previous moment. The horizontal coordinates of the row in the previous moment are represented by array a, which is an N×1-dimensional array, where N is the number of spectators in the row in the previous moment. The horizontal coordinates of the row in the current moment are represented by array b, which is an M×1-dimensional array, where M is the number of spectators in the row in the previous moment.
[0111] In one example, examples of array a and array b can be referred to as shown in Table 1.
[0112] Table 1
[0113] a 100 200 300 500 b 220 90 490 620
[0114] Among them, the horizontal coordinates in Table 1 are sorted according to the center height of the face detection frame. In one implementation method, taking the increasing horizontal coordinate values from left to right as an example, the height of the user on the left is taller, and the height of the user on the left is shorter. Then, when sorting according to the center height of the face detection frame, the user on the left is sorted later than the user on the right. At this time, the user on the left with a smaller horizontal coordinate will be located behind the user on the right with a larger horizontal coordinate, that is, 90 in Table 1 will be located to the right of 220.
[0115] Next, calculate the difference between the horizontal coordinates of each target user in the current keyframe and the horizontal coordinates of each target user in the previous keyframe. This means calculating the absolute difference between each element in the two arrays using the formula: D(i,j) = |a_i - b_j|, where i∈[1,N] and j∈[1,M].
[0116] For the specific calculation, using Table 2 as an example, take the number 100 in array a and calculate the difference between it and each value in array b. That is, calculate the difference between 100 and 220, 90, 490, and 620. Then, calculate the difference between the other values in array a, namely 200, 300, and 500, and each value in array b. The difference refers to the absolute difference. The specific calculation results are shown in Table 2 or stored in the constructed N×M dimensional distance matrix.
[0117] Table 2
[0118]
[0119] S32 . Match the target user in the key frame at the current moment with the target user in the key frame at the previous moment in ascending order of the difference values to obtain a user matching result.
[0120] Specifically, as shown in Table 2, the order of absolute differences from small to large is 10, 10, 20, 80, ... .
[0121] During the iterative matching process, the elements in array a are matched with the elements in array b in ascending order of absolute difference.
[0122] First, locate the index (i, j) of the element with the current minimum distance in the distance matrix. In practice, the minimum absolute difference is 10, which is the difference between 100 in array a and 90 in array b. Therefore, 100 in array a and 90 in array b can be combined to form a seat pair, namely (100, 90). At this point, (100, 90) is the first match in the user's matching results.
[0123] Then remove the matched a_i and b_j elements from the original arrays, that is, remove 100 from array a and 90 from array b. At this time, array a has 200, 300, and 500, and array b has 220, 490, and 620. And update the distance matrix according to the new arrays a and b.
[0124] Locate the position index (i, j) of the current minimum distance element from the new distance matrix. At this time, in the new distance matrix, the minimum absolute difference is 10, which is obtained by the difference between 500 in array a and 490 in array b. Then, 500 in array a and 490 in array b can be formed into a seat pair, that is, (500, 490).
[0125] Determine the positional relationship between the new seat pair (500, 490) and (100, 90) in the user matching results. Specifically, the positional relationship is that 500 in (500, 490) located in array a is greater than 100 in (100, 90) located in array a, and 490 in (500, 490) located in array b is greater than 90 in (100, 90) located in array b. The positional relationship verification passes, and the new seat pair (500, 490) is added to the user matching results. The seat pairs in the user matching results are sorted in ascending order according to a_i.
[0126] If the position relationship verification fails, the matching is completed and the iteration is exited.
[0127] Then, remove the matched a_i and b_j elements from the original array and dynamically update the distance matrix.
[0128] Locate the position index (i, j) of the current minimum distance element from the new distance matrix. At this time, in the new distance matrix, the minimum absolute difference is 20, which is obtained by subtracting 200 in array a and 220 in array b. Then, 200 in array a and 220 in array b can be combined into a seat pair, that is, (200, 220).
[0129] Determine the positional relationship between the new seat pair (200, 220) and the user matching results (100, 90) and (500, 490). The specific positional relationship is:
[0130] The seats (100, 90), (500, 490), and (200, 220) are sorted by size to (100, 90), (200, 220), and (500, 490). The elements 100, 200, and 500 in array a increase in size, and the elements 90, 220, and 190 in array b also increase in size. This indicates that for each seat pair (a_i, b_j), the sorting result for a_i is the same as the sorting result for b_j, indicating that the positional relationship verification has passed. At this point, the new seat pair (200, 220) is added to the user matching results, and the seat pairs in the user matching results are sorted in ascending order according to a_i. This means that the seat pairs are (100, 90), (200, 220), and (500, 490). The above steps are repeated until all elements in any array have been processed.
[0131] If the position relationship verification fails, the matching is completed and the iteration is exited.
[0132] Then, remove the matched a_i and b_j elements from the original array and dynamically update the distance matrix.
[0133] At this point, only 300 remains in array a and 620 remains in array b, forming (300, 620). A positional relationship check is then performed. For (100, 90), (200, 220), (300, 620), and (500, 490), elements 100, 200, 300, and 500 in array a increase in order, while elements 90, 220, 620, and 490 in array b do not follow this increasing order. Therefore, the positional relationship check fails, indicating that the user with horizontal coordinate 300 in array a and the user with horizontal coordinate 620 in array b are not the same person. Furthermore, the user with horizontal coordinate 620 was not present in the previous keyframe and may have returned temporarily. Matching is now complete, and the iteration is exited.
[0134] The user matching results obtained at this time are (100,90), (200,220) and (500,490).
[0135] S33 . For a user to be assigned that is not in the user matching result among the target users in the key frame at the current moment, determine the position information of the user to be assigned based on the horizontal coordinate of the user to be assigned.
[0136] Specifically, the target user in the key frame at the current moment is not located in the user matching result. The user with the horizontal coordinate of 620 is larger than the maximum value 500 in the array b in the user matching result. Therefore, the user with the horizontal coordinate of 620 should be inserted after (500, 490).
[0137] S34: Determine a user tracking result based on the user matching result and the location information of the user to be assigned.
[0138] Specifically, the final user tracking results are sorted from small to large according to the horizontal coordinates as follows:
[0139] (100,90), (200,220), (500,490) and 620.
[0140] After determining the user tracking results, each seat can be numbered, such as the body posture of the audience member of the seat ID, which can be divided into various action postures such as standing, swaying, and clapping.
[0141] In this embodiment, the seat matching method based on the minimum Euclidean distance matches the users at the previous moment and the current moment, and searches for the same user in the same seat number. By analyzing the posture changes of the user, the user's movement posture can be determined.
[0142] In one implementation, after determining the action information of the target user in the user tracking result, the seat number, play timestamp, emotion information, and action information of the target user may be stored in a target area.
[0143] In the specific implementation, the target area is the Redis cache, and the data list consisting of the target user's seat number, facial emotion, body posture, and playback timestamp is placed in the Redis cache to achieve real-time data storage.
[0144] In one implementation, referring to Figure 4 After storing the target user's seat number, playback timestamp, emotion information, and action information in the target area, the method further includes:
[0145] S41: Pull the latest data stored in the target area.
[0146] The data from the Redis cache is periodically pulled. Assume that the current data pull time, i.e., the pull timestamp, is timestamp. The timestamp of the latest data pulled from the Redis cache is data_lasttime. The period for periodic data pull can be the same as the duration of the aforementioned video segments, i.e., 2 seconds. Furthermore, the period for periodic data pull can be customized, such as 1 second.
[0147] S42: If the playback timestamp in the latest data is less than or equal to a preset time after the pulling time, extract the data of the specified time period stored in the target area at the preset time.
[0148] The specified time period is the time period between the pulling time and the preset time.
[0149] In specific implementation, if data_lasttime is less than or equal to timestamp, it means that the latest data has not been updated in the Redis cache. This may be due to data analysis delay, which causes the data to not be analyzed and stored in the Redis cache in time. In this case, the current batch of data is empty.
[0150] If data_lasttime is less than timestamp + the period of scheduled data pulling, where the period of scheduled data pulling can be 2 seconds or 1 second. Taking the period of scheduled data pulling as 2 seconds as an example, wait for timestamp+1-data_lasttime seconds. For example, if timestamp is 10 seconds, data_lasttime is 10.05 seconds, the period of scheduled data pulling is 1 second, timestamp+the period of scheduled data pulling is 11 seconds, and 10.05 seconds is less than 11 seconds, then wait for a certain period of time, which is specifically timestamp+the period of scheduled data pulling-data_lasttime, that is, wait for 0.5 seconds. At this time, the preset time is the time after waiting for 0.5 seconds. When the preset time is reached, the cached data between timestamp~timestamp+1 is pulled from redis so that the data can be pulled normally.
[0151] In one implementation, after extracting the data of the specified time period stored in the target area, the method further includes:
[0152] When the number of users in the latest data is greater than a preset number, current emotion distribution data of the user group in the latest data is determined, and based on the current emotion distribution data, emotion information filling operation is performed on users in the user group for whom no emotion information exists.
[0153] In specific implementation, in order to avoid the problem of incomplete display of users on the live broadcast screen due to the small amount of data of the detected emotion or action information, a data filling operation will be performed in this embodiment.
[0154] When filling in specific data, the number of users in the latest data is directly obtained. If the number of users is greater than a preset number, which can be 0.1 of the estimated audience size collected by the corresponding terminal, if the number of users in the latest data is greater than the preset number, it means that the latest data contains a large number of users. At this point, the current emotional distribution data of the user group in the latest data can be determined.
[0155] Among them, the current sentiment distribution data refers to the proportion of different emotions in the latest data.
[0156] For example, if there are 80 users in the latest data and 20 of them are happy, then the proportion of happy users is 0.25. The proportion of other emotions is determined in a similar way. The current emotion distribution data is used as the historical emotion distribution data.
[0157] Then, traverse the seat numbers. If a seat number does not exist in the latest data, it means that the key frame has not detected the emotions of the user in that seat. At this time, the historical emotion distribution data can be used to fill in the emotions of the users labeled with the seat numbers. For example, if there are 20 users whose emotions are not detected, then 4 of these 20 people are filled in as happy. The specific 4 people can be selected randomly or according to rules.
[0158] In one implementation, when the number of users in the latest data is not greater than a preset number, default emotion distribution data is used to perform an emotion information filling operation on users in the user group for whom emotion information does not exist.
[0159] Specifically, when the number of users in the latest data is not greater than the preset number, it means that the number of users collected in the latest data is small. At this time, the current emotion distribution data of the user group in the latest data may not be representative. The default emotion distribution data can be used to perform emotion information filling operations on users in the user group who do not have emotion information.
[0160] The default emotion distribution data may be configured according to actual conditions, such as all being calm, or half being happy and half being calm.
[0161] In this embodiment, data processing and backup operations are designed to fill in missing data, which can ensure the integrity of the data. After the data is completed, it is sent to the front end for use on the live broadcast page.
[0162] It should be noted that live broadcasts of concerts typically experience a delay, typically ranging from a few seconds to several minutes. In this embodiment, the live broadcast delay for the concert is 2.5 minutes, and the delay between capturing and analyzing the audience's footage is approximately 10 seconds. When the live broadcast page is sent, the 10-second delay is subtracted from the timestamp of the live broadcast data. When searching on the front end, the real-time time can be retrieved.
[0163] In summary, if Figure 5 As shown, in the embodiment of the present application, taking the above-mentioned terminal as a mobile phone as an example, the live video stream is uploaded to the cloud server through the collaborative collection of the audience seat images by multiple mobile phones. The cloud server tracks and analyzes the audience's expressions / actions. It is stored in the redis cache, integrated and supplemented data is sent to the front end, realizing real-time monitoring of the status of the audience seats, such as detecting that the 3rd row, 20 seats and 64 seats are standing up and cheering. It can be used for real-time collection and display of audience emotions in concerts or live broadcast scenes, enhancing the immersion and interactivity of the live broadcast scene. The embodiment of the present application has the following advantages:
[0164] Improved user experience: This embodiment of the application displays audience expressions and movements in real time, enhancing engagement and increasing click-through rates. When the system detects that more than 50% of the audience is "standing up and cheering," it triggers a resounding emotional response. The number of resounding responses after each singer's performance provides an objective indicator of the audience's enthusiasm for the song.
[0165] Reduced operating costs: This method can collect audience emotions and actions in real time. Using terminals instead of professional cameras and cloud servers reduces manual operating costs compared to the method where the director needs to manually find the audience and collect data. It creates a low-cost, high-precision, and highly real-time audience emotion and behavior analysis system, realizes multi-dimensional data collection and intelligent analysis, and becomes a key technological breakthrough in improving the live entertainment experience.
[0166] In addition, the multi-terminal embodiment of the present application collects audience images in real time, and the real-time analysis and data processing background processes the data before displaying it on the live page. This allows the audience to experience the live atmosphere in real time when watching the live performance, increasing the sense of immersion. At the same time, each singer can be evaluated with an objective emotional resonance score based on the emotional resonance evoked by the singer in each performance.
[0167] Based on the embodiment of the above-mentioned method for analyzing data of a user group, another embodiment of the present application provides a device for analyzing data of a user group, referring to Figure 6 , which may include:
[0168] An acquisition module 11 is configured to acquire key frames in a video segment currently acquired by at least one data acquisition terminal; each of the data acquisition terminals is configured to acquire video information of users in different areas of a user group;
[0169] The recognition module 12 is used to identify the facial information and body key point information corresponding to the same user in the key frame; the facial information at least includes emotional information;
[0170] A segmentation module 13 is configured to perform a location segmentation operation on the user's location based on the user's facial information to determine target users corresponding to the same location attribute;
[0171] A matching module 14 is configured to match the target user in the key frame at the current moment with the target user in the key frame at the previous moment based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, thereby obtaining a user tracking result;
[0172] The determination module 15 is configured to determine the action information of the target user in the user tracking result based on at least one of the facial information and the body key point information of the target user at different moments in the user tracking result.
[0173] In one implementation, the identification module 12 is specifically configured to:
[0174] The facial information and body key point information of the user in the key frame are identified; the facial information includes facial emotion information and facial key point information in the face detection frame, and based on the relative distance between the face detection frame and the specified key point in the body key point information, the facial information and the body key point information are matched to obtain facial information and body key point information corresponding to the same user.
[0175] In one implementation, the division module 13 is specifically configured to:
[0176] Based on the center height of the face detection frame in the facial information, the users are sorted to obtain a sorting result; based on the average height of the face detection frame and the mean of the center height of the face detection frame, the users in the sorting result are position-classified to determine target users belonging to the same position attribute; the position attribute includes a sorting number.
[0177] In one implementation, the matching module 14 includes:
[0178] a calculation submodule, configured to calculate a difference between a horizontal coordinate of each target user in the key frame at the current moment and a horizontal coordinate of each target user in the key frame at the previous moment;
[0179] A matching submodule, configured to match the target user in the key frame at the current moment with the target user in the key frame at the previous moment in ascending order of the difference values, to obtain a user matching result;
[0180] an information determination submodule, configured to determine, for a user to be assigned that is not located in the user matching result among target users in the key frame at the current moment, location information of the user to be assigned based on the horizontal coordinates of the user to be assigned;
[0181] The result determination submodule is configured to determine the user tracking result based on the user matching result and the location information of the user to be assigned.
[0182] In one implementation, the method further includes:
[0183] The cache module is used to store the seat number, play timestamp, emotion information and action information of the target user in the target area.
[0184] In one implementation, the method further includes:
[0185] A pulling module, used to pull the latest data stored in the target area;
[0186] The extraction module is used to extract the data of the specified time period stored in the target area at the preset time if the playback timestamp in the latest data is less than or equal to the preset time after the pulling time; the specified time period is the time period between the pulling time and the preset time.
[0187] In one implementation, the method further includes:
[0188] A first filling module is configured to determine current emotion distribution data of a user group in the latest data when the number of users in the latest data is greater than a preset number, and perform an emotion information filling operation on users in the user group for whom emotion information does not exist based on the current emotion distribution data;
[0189] The second filling module is configured to, when the number of users in the latest data is not greater than a preset number, use default emotion distribution data to perform an emotion information filling operation on users in the user group for whom no emotion information exists.
[0190] In this embodiment, at least one data acquisition terminal is used to collect video information of users in different areas of the user group, which can cover users in all areas and ensure the comprehensiveness of user analysis. Subsequently, the facial information and body key point information corresponding to the same user in the key frame are identified, and based on the facial information of the user, a position division operation is performed on the location of the user to determine the target user corresponding to the same position attribute. Based on the coordinates of the target user in the key frame of the current moment and the coordinates of the target user in the key frame of the previous moment, a matching operation is performed on the target user in the key frame of the current moment and the target user in the key frame of the previous moment to obtain the user tracking result. Based on at least one of the facial information and body key point information of the target user in the user tracking result at different moments, the action information of the target user in the user tracking result is determined. That is, in this application, the user's actions and emotions can be automatically identified based on the above process, avoiding the influence of human subjective factors on the recognition results and improving the accuracy of data recognition. In addition, during action recognition, the target user's facial information at different times and at least one of the body key point information in the user tracking results are determined. Since the body key points at different times are taken into account, that is, the user's posture at different times is taken into account, the accuracy of user action determination is improved.
[0191] It should be noted that, for the working process of each module and sub-module in this embodiment, please refer to the corresponding description in the above embodiment, which will not be repeated here.
[0192] An embodiment of the present application further provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0193] The memory is used to store computer programs;
[0194] The processor is used to execute the computer program so that the electronic device can implement the above-mentioned user group data analysis method.
[0195] refer to Figure 7 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0196] like Figure 7 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0197] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0198] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any user group data analysis method provided in the embodiment of the present application.
[0199] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any user group data analysis method provided in the embodiment of the present application.
[0200] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for analyzing user group data, characterized in that: include: Obtaining key frames in video segments collected by at least one data collection terminal at a current moment; Each of the data collection terminals is used to collect video information of users in different areas of the user group; Identifying facial information and body key point information corresponding to the same user in the key frame; the facial information at least includes emotional information; Based on the facial information of the user, performing a location division operation on the user's location to determine target users corresponding to the same location attribute; Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result; Based on at least one of facial information and body key point information of the target user in the user tracking result at different moments, action information of the target user in the user tracking result is determined.
2. The method for analyzing user group data according to claim 1, characterized in that: Identifying facial information and body key point information corresponding to the same user in the key frame, including: Identify the user's facial information and body key point information in the key frame; the facial information includes facial emotion information and facial key point information in the face detection frame; Based on the relative distance between the face detection frame and the designated key point in the human body key point information, a matching operation is performed on the face information and the human body key point information to obtain the face information and the human body key point information corresponding to the same user.
3. The method for analyzing user group data according to claim 1, wherein: Based on the facial information of the user, a location segmentation operation is performed on the location of the user to determine target users corresponding to the same location attribute, including: performing a sorting operation on the users based on the center heights of the face detection frames in the face information to obtain a sorting result; Based on the average height of the face detection frame and the average center height of the face detection frame, a position classification operation is performed on the users in the sorting result to determine target users belonging to the same position attribute; the position attribute includes a sorting number.
4. The method for analyzing user group data according to claim 1, wherein: Based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, a matching operation is performed on the target user in the key frame at the current moment and the target user in the key frame at the previous moment to obtain a user tracking result, including: Calculating the difference between the horizontal coordinate of each target user in the key frame at the current moment and the horizontal coordinate of each target user in the key frame at the previous moment; Matching the target user in the key frame at the current moment with the target user in the key frame at the previous moment in ascending order of the difference values to obtain a user matching result; For a user to be assigned that is not in the user matching result among the target users in the key frame at the current moment, determining the position information of the user to be assigned based on the horizontal coordinate of the user to be assigned; A user tracking result is determined based on the user matching result and the location information of the user to be assigned.
5. The method for analyzing user group data according to claim 1, wherein: After determining the action information of the target user in the user tracking result, the method further includes: The seat number, play timestamp, emotion information and action information of the target user are stored in the target area.
6. The method for analyzing user group data according to claim 5, characterized in that: After storing the target user's seat number, play timestamp, emotion information, and action information in the target area, the method further includes: Pulling the latest data stored in the target area; If the playback timestamp in the latest data is less than or equal to a preset time after the pulling time, at the preset time, data of a specified time period stored in the target area is extracted; the specified time period is the time period between the pulling time and the preset time.
7. The method for analyzing user group data according to claim 6, wherein: After extracting the data of the specified time period stored in the target area, the method further includes: When the number of users in the latest data is greater than a preset number, determining current emotion distribution data of the user group in the latest data; Based on the current emotion distribution data, performing an emotion information filling operation on users in the user group for whom no emotion information exists; When the number of users in the latest data is not greater than a preset number, default emotion distribution data is used to perform an emotion information filling operation on users in the user group for whom emotion information does not exist.
8. A data analysis device for a user group, characterized in that: include: An acquisition module, configured to acquire key frames in a video segment acquired by at least one data acquisition terminal at a current moment; Each of the data collection terminals is used to collect video information of users in different areas of the user group; A recognition module, configured to identify facial information and body key point information corresponding to the same user in the key frame; The facial information at least includes emotional information; a segmentation module, configured to perform a location segmentation operation on the user's location based on the user's facial information, so as to determine target users corresponding to the same location attribute; a matching module, configured to match the target user in the key frame at the current moment with the target user in the key frame at the previous moment based on the coordinates of the target user in the key frame at the current moment and the coordinates of the target user in the key frame at the previous moment, to obtain a user tracking result; The determination module is configured to determine the action information of the target user in the user tracking result based on at least one of the facial information and the body key point information of the target user at different moments in the user tracking result.
9. An electronic device, characterized in that: comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program so that the electronic device can implement the user group data analysis method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the user group data analysis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Luggage case back following method and a luggage case side following method based on behavior prediction
CN114022929A
Target tracking method and device, electronic equipment and computer readable storage medium
CN115760905A
Emotion recognition method based on video analysis technology and upper limb pose description
CN118411745A
System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform
US20220121884A1