Personalized live broadcast method and device, electronic equipment and computer storage medium
By generating a virtual seating chart through real-time acquisition of on-site image data, and performing personalized image processing on user terminals, the problems of difficulty in audience focusing and high equipment costs in traditional live streaming solutions are solved, achieving a personalized live streaming experience and load optimization.
Patent Information
- Application Number
- CN202511311706.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-12
AI Technical Summary
Traditional live streaming solutions cannot accurately target online viewers, lacking focus and immersion. Furthermore, they are costly in terms of equipment and data processing in large audience scenarios, making them difficult to apply widely.
By deploying a limited number of camera devices, real-time image data is collected from the site, generating a virtual seating map. User terminals perform real-time image processing based on target cropping information to generate terminal display screens, reducing the server-side streaming pressure.
It enhances the personalized experience for users during live streaming interactions, significantly reduces server-side streaming pressure, and optimizes live streaming load under large-scale concurrency.
Smart Images

Figure CN121126015A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the online live broadcast technical field, and particularly relates to a personalized live broadcast method and device, electronic equipment and computer storage medium. BACKGROUND
[0002] With the increasing richness of live event live broadcast scenes, viewers have higher requirements for the personalization and interactivity of live broadcast content.
[0003] At present, the traditional live broadcast scheme usually covers the panorama through multiple positions, but online viewers cannot accurately focus on a specific viewer or seat area, and lack of pertinence and immersion. At the same time, if an independent camera device is set for each viewer for live broadcast collection, although local picture acquisition can be realized, the device cost, wiring and data processing pressure are extremely high, and it is difficult to be widely applied in large-scale viewer scenes. SUMMARY
[0004] Therefore, the present application provides a personalized live broadcast method, device, electronic equipment and computer storage medium, which effectively improves the experience of users in the live broadcast interaction process, significantly reduces the server-side streaming pressure, and realizes live broadcast load optimization under large-scale concurrency.
[0005] The first aspect of the present application provides a personalized live broadcast method, comprising:
[0006] real-time collection of live image data based on a camera device;
[0007] generating a virtual seat mapping table according to the live image data;
[0008] receiving a virtual seat selection request of a user; wherein the virtual seat selection request includes a target virtual seat selected by the user and user terminal information;
[0009] retrieving target cropping information in the virtual seat mapping table according to the identifier of the target virtual seat;
[0010] downloading the target cropping information to a user terminal; wherein the user terminal performs real-time image processing on the current live frame according to the target cropping information to generate a terminal display picture.
[0011] Optionally, the generating a virtual seat mapping table according to the live image data comprises:
[0012] determining an effective seat in the live image data based on a preset seat structure template;
[0013] for each of the effective seats, recording the center point coordinates of the effective seat, and assigning a unique number to the effective seat;
[0014] convert the coordinates of the effective agent into virtual coordinates in a virtual interface based on the center point coordinates of the effective agent;
[0015] generate structured data of the effective agent according to the unique number of the effective agent, the coordinates of the effective agent, the virtual coordinates of the effective agent, the camera number of the region to which the effective agent belongs, and the image cropping region;
[0016] generate a virtual agent mapping table based on the structured data of all the effective agents.
[0017] Optionally, after determining the effective agent in the live image data based on the preset agent structure template, the method further comprises:
[0018] perform human body recognition on the region where the effective agent is located to obtain a human body recognition result, wherein the human body recognition result is whether there is a human being in the region of the effective agent;
[0019] based on the human body recognition result, perform personnel information labeling on the effective agent in a virtual interface.
[0020] Optionally, after generating the virtual agent mapping table according to the live image data, the method further comprises:
[0021] perform consistency detection on the virtual agent mapping table to obtain a consistency detection result;
[0022] for each abnormal region in the consistency detection result, correct the virtual coordinates in the abnormal region.
[0023] Optionally, the position and angle of each camera device are determined according to a floor plan of the site and a view angle parameter of the camera device, wherein the floor plan of the site includes agent distribution information and stage position, and the view angle parameter of the camera device includes focal length, horizontal view angle, and vertical view angle.
[0024] Optionally, the personalized live streaming method further comprises:
[0025] if the region to which the agent belongs is covered by multiple camera devices at the same time, fuse the images collected by the multiple camera devices to obtain a fused image.
[0026] Optionally, the personalized live streaming method further comprises:
[0027] determine a priority score of the camera device according to the number of times the target virtual agent is selected, the number of agents covered by the camera device, and the set of image cropping regions covered by the camera device, and dynamically allocate image processing resources and transmission bandwidth based on the priority score of the camera device.
[0028] The second aspect of the present application provides a personalized live streaming device, comprising:
[0029] An acquisition unit is configured to acquire live image data in real time based on a camera device;
[0030] A generation unit is configured to generate a virtual seat mapping table according to the live image data;
[0031] A receiving unit is configured to receive a virtual seat selection request of a user; wherein the virtual seat selection request comprises a target virtual seat selected by the user and user terminal information;
[0032] A retrieving unit is configured to retrieve target cropping information from the virtual seat mapping table according to an identifier of the target virtual seat;
[0033] An issuing unit is configured to issue the target cropping information to a user terminal; wherein the user terminal performs real-time image processing on a current live streaming frame according to the target cropping information to generate a terminal display screen.
[0034] Optionally, the generation unit comprises:
[0035] An effective seat determination unit is configured to determine an effective seat in the live image data based on a preset seat structure template;
[0036] A record allocation unit is configured to record a center point coordinate of each effective seat and allocate a unique number to the effective seat;
[0037] A coordinate conversion unit is configured to convert a coordinate of the effective seat into a virtual coordinate in a virtual interface based on the center point coordinate of the effective seat;
[0038] A structured data generation unit is configured to generate structured data of the effective seat according to a unique number of the effective seat, a coordinate of the effective seat, a virtual coordinate of the effective seat, a camera number of a region to which the effective seat belongs, and an image cropping region;
[0039] A virtual seat mapping table generation unit is configured to generate a virtual seat mapping table based on structured data of all the effective seats.
[0040] Optionally, the personalized live streaming device further comprises:
[0041] A human form recognition unit is configured to perform human form recognition on a region where the effective seat is located to obtain a human form recognition result; wherein the human form recognition result indicates whether there is a person in the region of the effective seat;
[0042] A labeling unit is configured to label personnel information of the effective seat in a virtual interface based on the human form recognition result.
[0043] Optionally, the personalized live broadcast device further comprises:
[0044] a consistency detection unit configured to perform consistency detection on the virtual host mapping table to obtain a consistency detection result;
[0045] a correction unit configured to correct the virtual coordinates in each abnormal area in the consistency detection result.
[0046] Optionally, the position and angle of each camera device are determined according to a floor plan and a view angle parameter of the camera device; the floor plan comprises host distribution information and stage position; and the view angle parameter comprises focal length, horizontal view angle and vertical view angle.
[0047] Optionally, the personalized live broadcast device further comprises:
[0048] a fusion unit configured to fuse images collected by multiple camera devices if the host belongs to an area covered by the multiple camera devices to obtain a fused image.
[0049] Optionally, the personalized live broadcast device further comprises:
[0050] a dynamic allocation unit configured to determine a priority score of the camera device according to the number of times the target virtual host is selected, the number of hosts covered by the camera device and the set of image cropping areas covered by the camera device, and dynamically allocate image processing resources and transmission bandwidth based on the priority score of the camera device.
[0051] The third aspect of the present application provides an electronic device comprising:
[0052] one or more processors;
[0053] a storage device having one or more programs stored thereon;
[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement the personalized live broadcast method according to any one of the first aspect.
[0055] The fourth aspect of the present application provides a computer storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the personalized live broadcast method according to any one of the first aspect.
[0056] From the above scheme, the application provides a personalized live broadcast method, device, electronic equipment and computer storage medium, through deploying a limited number of camera equipment, real-time collection of live image data is realized, and one-to-one mapping from actual image to online virtual seat is realized, a virtual seat mapping table is obtained, so that after the user clicks the virtual seat, the system no longer generates an independent live broadcast stream, but only issues the target cutting information corresponding to the target virtual seat selected by the user in the virtual seat mapping table to the user terminal, and the user terminal performs real-time image processing on the current live frame according to the target cutting information to generate a terminal display picture; effectively improving the experience of the user in the live broadcast interaction process, while significantly reducing the server-side stream pushing pressure, realizing live broadcast load optimization under large-scale concurrency. BRIEF DESCRIPTION OF DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute a part of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on the provided drawings.
[0058] Figure 1 A specific flowchart of a personalized live broadcast method provided by the embodiment of the present application is provided.
[0059] Figure 2 A specific flowchart of a method for generating a virtual seat mapping table provided by another embodiment of the present application is provided.
[0060] Figure 3 A schematic diagram of a virtual interface provided by another embodiment of the present application is provided.
[0061] Figure 4 A schematic diagram of a personalized live broadcast device provided by another embodiment of the present application is provided.
[0062] Figure 5 A schematic diagram of an electronic device for realizing a personalized live broadcast method provided by another embodiment of the present application is provided. DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0064] The term "include" and variations thereof, as used in this document, mean "to include, without limitation." The term "based on" means "based at least in part on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms shall be construed accordingly.
[0065] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.
[0066] It should be noted that the "first", "second", and the like concepts mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence of the functions performed by these devices, modules or units.
[0067] It should be noted that the modification of "one" or "multiple" in the present application is illustrative and not limiting, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0068] The embodiment of the present application provides a personalized live streaming method, as shown in the figure, which specifically comprises the following steps: Figure 1 As shown in the figure, specifically comprising the following steps:
[0069] S101, real-time collection of live image data based on a camera device.
[0070] Optionally, in another embodiment of the present application, the position and angle of each camera device are determined according to the floor plan and the angle of view parameters of the camera device.
[0071] The floor plan includes the seat distribution information and the stage position; the angle of view parameters of the camera device include the focal length, the horizontal angle of view and the vertical angle of view.
[0072] In the specific implementation process of the present application, the way to determine the position and angle of the camera device can be, but is not limited to, establishing a two-dimensional plane coordinate system according to the floor plan, parameterizing the possible installation position and height of the camera, and then importing the angle of view parameters (focal length, horizontal angle of view, vertical angle of view) of the camera in the modeling software; then, the coverage range of each camera point is simulated, and the visible area is marked on the graph; by adjusting the position and angle of the camera, each point covers as many seats as possible, so as to ensure that all seat areas are completely covered by at least one camera, which is not limited here.
[0073] In the specific implementation of the present application, 15%-30% of the overlapping area at the boundary can also be reserved for subsequent image fusion and redundancy correction.
[0074] In the specific implementation of the present application, the image frames collected by each camera can also be uniformly spatially calibrated, including but not limited to perspective correction, picture center unification, and image partition numbering.
[0075] In the specific implementation of the present application, the specific implementation of spatial calibration can be but is not limited to arranging multiple calibration points (positions known and fixed) in the site; then, each camera shoots these calibration points, and the mapping matrix from the camera coordinate system to the global coordinate system is calculated through multi-point perspective transformation; then, according to the mapping matrix, the perspective distortion is corrected, and the center of the image is aligned to the global reference point; finally, the corrected image is divided into multiple logical regions (RegionID), each region corresponds to the seat number of the global seat map;
[0076] It can be understood that the calibration process is performed at the initial installation and when the camera position changes.
[0077] In the actual application of the present application, the field of view of each camera is divided into several logical grid units, and each unit records its coordinate position and number in the global seat map, and the following calibration table is generated :
[0078] ;
[0079] Among them, is the camera equipment identifier, is the image cropping area, and Global Coord is the global seat map coordinate system. Through the calibration table, the subsequent image synthesis, seat recognition, and region cropping have clear corresponding relationship.
[0080] Optionally, in another embodiment of the present application, if the seat belongs to the region covered by multiple camera equipment, the images collected by multiple camera equipment are fused to obtain the fused image.
[0081] The specific fusion method can use but is not limited to the weighted average method, so as to improve the image quality and recognition accuracy. In the actual application process, the following calculation formula can be used to calculate the fused image :
[0082] ;
[0083] Among them, represents the weight value of the pixel point in the two images, which is usually set according to the factors such as definition and center offset distance. These represent image frames captured by the two cameras in the overlapping area; It represents the two-dimensional coordinate position of a pixel in the image frame.
[0084] Of course, multiple images can be compared at the pixel level to select the version with higher clarity and no obstruction as the main image; and when a certain image fails (such as obstruction or camera failure), the present invention can automatically switch to the backup image to ensure the continuity of the image, which is not limited here.
[0085] In the practical application of this invention, video stream synchronization and timing calibration will also be performed. Specifically, a timestamp synchronization and frame order unification strategy can be adopted to ensure that subsequent images from multiple camera sources can be accurately cropped and mapped at the frame level. The keyframe alignment formula is as follows:
[0086] ;
[0087] in, For reference timeline, For the first Street camera frame timestamps, time differences between all video frames Keep it within the set threshold (e.g., 40ms).
[0088] S102. Generate a virtual seat mapping table based on the on-site image data.
[0089] In the specific implementation of this invention, before generating the virtual seat mapping table based on the on-site images, the on-site image data can also be standardized, including but not limited to distortion correction, brightness equalization, noise removal, etc., which are not limited here, in order to generate standard image data with a unified input format.
[0090] Optionally, in another embodiment of the present invention, one implementation of step S102 is as follows: Figure 2 As shown, it includes:
[0091] S201. Determine the valid seats in the on-site image data based on the preset seat structure template.
[0092] In this context, "valid seating" refers to the location of a physical audience seat in the venue layout, rather than the background, aisle, or other non-seat areas.
[0093] In the specific implementation of this invention, the matching area in the field image can be searched by sliding the preset seat structure template, but is not limited to, to determine the valid seats in the field image data. This is not limited here.
[0094] The template sliding matching method can be as follows:
[0095] ; represents the matching score of the image starting from the position , is the input image (live image data), is the seat template (preset seat structure template). m and n represent the row coordinate and column coordinate positions of the seat template respectively. m and n are used to traverse the two-dimensional pixel array of the template to calculate the matching score of the template at a certain position of the image by point-by-point multiplication and summation with the image area after standardization processing. When the matching score exceeds the set threshold, the area is identified as a valid seat.
[0096] S202, for each valid seat, record the center point coordinates of the valid seat, and the valid seat is assigned a unique number.
[0097] wherein the center point coordinates are pixel coordinate positions in the camera shot picture, and the center point coordinates are used to accurately calculate the position of the seat in the camera view, and assist in establishing the mapping relationship between the physical seat and the virtual seat. The unique number (UID) is assigned based on the logical position of the seat in the global layout (such as the row and column number, the camera number) and the identified session information once. The UID is used as the logical identifier of the valid seat, and the center point coordinates are used for geometric mapping calculation.
[0098] S203, based on the center point coordinates of the valid seat, convert the coordinates of the valid seat into virtual coordinates in the virtual interface.
[0099] In the implementation process of the present application, the corresponding virtual seat layer is drawn on the virtual interactive interface to obtain the virtual interface; through the spatial mapping relationship between the preset seat structure template and the position of the valid seat in the live image data, a mapping function from the actual image coordinates (coordinates of the valid seat) to the virtual seat coordinates (virtual coordinates in the virtual interface) is constructed:
[0100] ;
[0101] wherein, represents the coordinates of the kth valid seat in the live image, is the mapping position of the valid seat k in the virtual interface, that is, the virtual coordinates of the valid seat.
[0102] The specific mapping operation can but is not limited to the following: mapping the pixel coordinate system of the camera shooting picture to the same reference system (such as normalized to 0~1 relative coordinates) as the coordinate system of the virtual interface agent layout; then, according to the relative position of the center point coordinate in the shooting picture, matching to the virtual agent point in the corresponding position in the virtual interface; and then, through the proportional scaling and offset calculation, converting the physical agent center point position to the virtual interface position.
[0103] That is, the determination of the mapping relationship is based on the initial calibration (the angle of view when the camera is installed and the agent layout of the site) and the agent detection result in the real-time image.
[0104] S204, generating the structured data of the valid agent according to the unique number of the valid agent, the coordinate of the valid agent, the virtual coordinate of the valid agent, the camera number of the region to which the valid agent belongs, and the image cropping region.
[0105] In the specific implementation process of the application, the structured data of the valid agent k can but is not limited to the following format, which is not limited here:
[0106] ;
[0107] Among them, represents the ID of the valid agent k, represents the camera number of the region to which the valid agent belongs, represents the image cropping region, which is determined by the agent center point coordinate and the effective region size around it. The specific determination method can but is not limited to defining a rectangular region with the agent center point as the center, combining the recognized agent width and height parameters; then, the region boundary is appropriately expanded outward to contain the complete range of audience body movements; finally, the region coordinate is recorded in the form of pixel coordinates of the camera picture, ensuring that the subsequent player end can accurately crop the region.
[0108] In the actual application process of the application, The size and proportion of the image cropping region can be parameterized and configured according to the resolution of the camera, the shooting distance, and the proportion of the audience in the picture during the initial deployment, and can also be dynamically fine-tuned through the recognition algorithm during operation.
[0109] S205, generating a virtual agent mapping table based on the structured data of all valid agents.
[0110] In the specific implementation process of the application, the format of the virtual agent mapping table can but is not limited to the following format, which is not limited here:
[0111] ;
[0112] Among them, The virtual seat mapping table is used to represent the virtual seat mapping table.
[0113] Optionally, in another embodiment of the present application, the personalized live broadcast method further comprises the following steps:
[0114] The human body recognition result is obtained by performing human body recognition on the area where the valid seat is located.
[0115] The human body recognition result is whether there is a person in the area where the valid seat is located.
[0116] In the specific implementation process of the present application, the manner of performing human body recognition on the area where the valid seat is located can be, but is not limited to, a target detection algorithm based on a convolutional neural network, etc., which is not limited here.
[0117] The present application only labels the seats where actual people are seated in the mapping and subsequent personalized playing by performing human body recognition on the area where the valid seat is located, avoiding meaningless mapping and interaction generation for empty seats. It can effectively improve resource utilization efficiency and reduce the possibility of subsequent player end cutting and displaying no content area.
[0118] As shown in Figure 3 Fig. 1 is a schematic diagram of a virtual interface provided by the present application, which clearly shows which seats are valid, which valid seats have people seated, and which valid seats have no people seated.
[0119] Optionally, in another embodiment of the present application, the present application provides a complete mapping correction and management mechanism based on the generation of the virtual seat mapping table, which is used to improve the mapping accuracy, cope with dynamic changes on site, and ensure that the virtual seat interface and the on-site audience picture always maintain a stable and consistent correspondence relationship. The personalized live broadcast method further comprises the following steps:
[0120] The consistency of the virtual seat mapping table is detected to obtain a consistency detection result; and the virtual coordinates in the abnormal area are corrected for each abnormal area in the consistency detection result.
[0121] In the actual application process of the present application, the overall mapping result can be detected for consistency by using, but not limited to, the seat structure continuity rule to judge whether there are local misplacement, seat number conflict or image drift problems. The consistency detection indicators include the distance between adjacent seats, the rationality of the number sequence, the mapping density, etc.
[0122] Specifically, according to the virtual coordinates of all identified agents in the mapping result, the row spacing and column spacing between adjacent agents can be calculated; then, the spacings are compared with the preset standard spacing range (such as a deviation of ±5%); then, it is checked whether the order of the row and column numbers is continuous and without repetition; it is judged whether the number of agents in the same row (column) is consistent with the preset of the venue layout; if it is found that the spacing between adjacent agents is abnormal, the numbering is dislocated or the number is inconsistent, the region is marked as a consistency abnormal region, and a local correction process is entered, so as to ensure that accurate mapping can be provided in multiple scenes and complex environments.
[0123] For the detected problem area, the application adopts a local rematching strategy to reanalyze the image information of the area, and performs interpolation correction with the surrounding known mapping result. The specific method can be but is not limited to a weighted local interpolation method, and the corrected mapping point coordinates are calculated as follows:
[0124] ;
[0125] Wherein, is the corrected virtual coordinate, is the surrounding confirmed virtual coordinates, is the weight based on the proximity of agents.
[0126] In the specific implementation process of the application, a background interface can also be provided to display the mapping relationship between the live image and the virtual agent layer. The user can fine-tune the agent corresponding relationship through dragging or clicking, and the system synchronously records the adjustment result and updates the mapping database. This function is particularly suitable for the following scenarios: the camera position is offset but has not been recalibrated; some audience temporarily changes seats or stands to block; the virtual interface needs to be fine-tuned for custom layout.
[0127] In the specific implementation process of the application, a mapping state detection task can also be periodically performed during live operation, such as identifying obvious personnel movement, scene obstruction or camera angle change in the image, which will automatically trigger local area re-identification and mapping update to ensure continuous and effective data.
[0128] In the specific implementation process of the application, all mapping relationships can also be written into a database in a structured data form, with a time stamp and a version number, supporting fast switching and calling of multiple sessions and multiple scenes. The data structure may be as follows:
[0129] ;
[0130] Wherein, StatusFlag is used to mark whether the mapping is manually adjusted, whether it is effective for the current version, and the like. represents an image cropping area, represents a version identification number. represents the coordinates of the kth valid seat in the live image, represents the mapping position of the kth valid seat in the virtual interface. represents the mapping position of the kth valid seat in the virtual interface.
[0131] S103, receiving a virtual seat selection request of a user.
[0132] The virtual seat selection request includes a target virtual seat selected by the user and user terminal information.
[0133] S104, retrieving target cropping information in the virtual seat mapping table according to the identification of the target virtual seat.
[0134] The target cropping information includes but is not limited to camera ID, image cropping region (ROI), image stretching scale (Scale), and optional perspective transformation matrix (Warp), etc. without limitation.
[0135] The following is an example of the identification of the target virtual seat k corresponding to .
[0136] .
[0137] wherein, is the image cropping region of the target virtual seat k selected by the user in the live stream; is the scale factor by which the image cropping region k needs to be enlarged or reduced; is the transformation parameter required for correction if there is perspective distortion in the image cropping region k.
[0138] In the initial mapping stage, the camera shooting distance, resolution, and the proportion of the audience in the picture are automatically calculated and set to ensure clear picture and moderate figure proportion. In the system settings, users can fine-tune this proportion on the player side, so as to enlarge or reduce the view according to personal preference, but the default value comes from the system automatic calculation result.
[0139] When initially deployed, it is automatically generated through a camera calibration process. This process can include but is not limited to: shooting calibration points at known positions on the shooting site (such as stage edges, seat markers); then calculating the perspective transformation matrix according to the physical coordinates and image coordinates of the calibration points; finally, the parameter can be recalculated according to the change of camera angle or site adjustment in running.
[0140] S105, sending the target cropping information to the user terminal.
[0141] Wherein, the user terminal carries out real-time image processing on the current live frame according to the target cropping information to generate a terminal display picture.
[0142] Specifically, the terminal player extracts a specified region from the current live frame after receiving the target cropping information, and performs image processing according to the proportion and the transformation matrix to generate a terminal display picture.
[0143] ;
[0144] Wherein, is an original live frame; is a terminal display picture; warp_resize represents an image function of uniform processing of cropping, scaling, perspective transformation and the like, that is, target cropping information or generation according to target cropping information.
[0145] It should be emphasized that the present application performs picture cropping, scaling and transformation processing locally by the terminal player, without the need to request a new live stream or a cloud video cropping service, but only based on the existing main live stream for localized picture processing. Thus, while realizing personalized live experience of user's perspective, the number of server stream pushing and transcoding cost are greatly reduced, meeting the business needs of high concurrency and low resource occupation.
[0146] In the specific implementation process of the present application, the terminal player can also support but is not limited to main picture replacement, floating window, picture-in-picture and other display modes, improving the flexibility of interaction, which is not limited here.
[0147] In the specific implementation process of the present application, in order to improve processing efficiency, the player locally caches frequently used cropping parameters and preprocessed image frames to avoid repeated calculation. The system sets a priority caching strategy for hot areas and high-click UIDs, and uses browser GPU or mobile hardware acceleration interface to perform image cropping and rendering, ensuring smooth playback experience.
[0148] It should be noted that since all users share one main live stream, only the local content of the picture is processed as needed, so the system has high scalability and can support hundreds of users to watch different areas at the same time without increasing the server stream pushing load, greatly enhancing the concurrent carrying capacity and maintainability of the system.
[0149] Optionally, in another embodiment of the present application, in order to avoid overload, an embodiment of the personalized live method of the present application further comprises:
[0150] The priority score of the camera equipment is determined according to the number of times the target virtual seat is selected, the number of seats covered by the camera equipment and the set of image cropping regions covered by the camera equipment, and the image processing resources and transmission bandwidth are dynamically allocated based on the priority score of the camera equipment.
[0151] Specifically, the priority score of the camera device can be determined by using the following formula:
[0152] ;
[0153] wherein, is the priority score of the camera device at present; is the number of all seats covered by the camera device; is the number of times that the user clicks the virtual seat corresponding to the UID within a certain time window; is the set of all ROIs (image cropping regions) covered by the camera device.
[0154] It can be understood that the camera device with a higher priority score can obtain more image processing resources and transmission bandwidth.
[0155] As can be seen from the above scheme, the application provides a personalized live streaming method, live image data is collected in real time by deploying a limited number of camera devices, and one-to-one mapping from actual image to online virtual seat is realized to obtain a virtual seat mapping table, so that after the user clicks the virtual seat, the system no longer generates an independent live streaming, but only issues the target cropping information corresponding to the target virtual seat selected by the user in the virtual seat mapping table to the user terminal, and the user terminal performs real-time image processing on the current live frame according to the target cropping information to generate a terminal display picture; the experience of the user in the live streaming interaction process is effectively improved, the server-side streaming pressure is significantly reduced, and live streaming load optimization under large-scale concurrency is realized.
[0156] Another embodiment of the application provides a personalized live streaming device, as shown in Figure 4 specifically comprising:
[0157] The acquisition unit 401 is configured to collect live image data in real time based on the camera device.
[0158] Optionally, in another embodiment of the application, the position and angle of each camera device are determined according to the floor plan and the angle of view parameter of the camera device.
[0159] The floor plan includes seat distribution information and stage position, and the angle of view parameter of the camera device includes focal length, horizontal angle of view and vertical angle of view.
[0160] The specific working process of the units disclosed in the above embodiments of the application can be referred to the corresponding method embodiments, which will not be described here again.
[0161] The generation unit 402 is configured to generate a virtual seat mapping table according to the live image data.
[0162] Optionally, in another embodiment of the application, an implementation of the generating unit 402 comprises:
[0163] An effective seat determining unit is configured to determine effective seats in the live image data based on a preset seat structure template.
[0164] A record assigning unit is configured to record, for each effective seat, a center point coordinate of the effective seat and assign a unique number to the effective seat.
[0165] A coordinate converting unit is configured to convert the coordinates of the effective seat into virtual coordinates in the virtual interface based on the center point coordinate of the effective seat.
[0166] A structured data generating unit is configured to generate structured data of the effective seat according to the unique number of the effective seat, the coordinates of the effective seat, the virtual coordinates of the effective seat, a camera number of a region to which the effective seat belongs, and an image cropping region.
[0167] A virtual seat mapping table generating unit is configured to generate a virtual seat mapping table based on the structured data of all the effective seats.
[0168] The specific working processes of the units disclosed in the above embodiments of the application can be found in the corresponding method embodiments, and thus will not be described here again. Figure 2
[0169] Optionally, in another embodiment of the application, an implementation of the personalized live streaming device further comprises:
[0170] A human form recognizing unit is configured to perform human form recognition on a region where the effective seat is located to obtain a human form recognition result.
[0171] The human form recognition result is whether there is a human being in the region of the effective seat.
[0172] A marking unit is configured to mark personnel information of the effective seat in the virtual interface based on the human form recognition result.
[0173] The specific working processes of the units disclosed in the above embodiments of the application can be found in the corresponding method embodiments, and thus will not be described here again.
[0174] Optionally, in another embodiment of the application, an implementation of the personalized live streaming device further comprises:
[0175] A consistency detecting unit is configured to perform consistency detection on the virtual seat mapping table to obtain a consistency detection result.
[0176] A correcting unit is configured to correct the virtual coordinates in each abnormal region in the consistency detection result.
[0177] The specific working process of the units disclosed in the above embodiments of the application can be seen from the corresponding method embodiment contents, which will not be repeated here.
[0178] Optionally, in another embodiment of the application, an embodiment of the personalized live streaming device further comprises:
[0179] The fusion unit is configured to fuse images collected by the multiple camera devices to obtain a fused image when the region to which the agent belongs is simultaneously covered by the multiple camera devices.
[0180] The specific working process of the units disclosed in the above embodiments of the application can be seen from the corresponding method embodiment contents, which will not be repeated here.
[0181] The receiving unit 403 is configured to receive a virtual agent selection request of a user.
[0182] The virtual agent selection request includes a target virtual agent selected by the user and user terminal information.
[0183] The searching unit 404 is configured to search for target cropping information in the virtual agent mapping table according to the identifier of the target virtual agent.
[0184] The issuing unit 405 is configured to issue the target cropping information to the user terminal.
[0185] The user terminal performs real-time image processing on a current live streaming frame according to the target cropping information to generate a terminal display picture.
[0186] The specific working process of the units disclosed in the above embodiments of the application can be seen from the corresponding method embodiment contents, which will not be repeated here. Figure 1
[0187] Optionally, in another embodiment of the application, an embodiment of the personalized live streaming device further comprises:
[0188] The dynamic allocation unit is configured to determine a priority score of the camera device according to the number of times the target virtual agent is selected, the number of agents covered by the camera device, and the set of image cropping regions covered by the camera device, and dynamically allocate image processing resources and transmission bandwidth based on the priority score of the camera device.
[0189] The specific working process of the units disclosed in the above embodiments of the application can be seen from the corresponding method embodiment contents, which will not be repeated here.
[0190] From the above scheme, the application provides a personalized live broadcast device, through deploying a limited number of camera equipment, real-time collection of live image data is realized, and one-to-one mapping from actual image to online virtual seat is realized, a virtual seat mapping table is obtained, so that after a user clicks a virtual seat, the system no longer generates an independent live broadcast stream, but only issues target cutting information corresponding to a target virtual seat selected by the user in the virtual seat mapping table to a user terminal, and the user terminal performs real-time image processing on a current live broadcast frame according to the target cutting information to generate a terminal display picture, thereby effectively improving the experience of the user in the live broadcast interaction process, significantly reducing the server-side stream pushing pressure, and realizing live broadcast load optimization under large-scale concurrency.
[0191] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0192] Another embodiment of the application provides an electronic device, such as Figure 5 as shown, comprising:
[0193] One or more processors 501.
[0194] Storage device 502, one or more programs are stored on the storage device 502.
[0195] When the one or more programs are executed by the one or more processors 501, the one or more processors 501 implement the personalized live broadcast method as described in the above embodiments.
[0196] Another embodiment of the application provides a computer storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the personalized live broadcast method as described in the above embodiments.
[0197] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0198] It is noted that the computer-readable medium described above in the present application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a computer-readable storage medium in a baseband or propagated as a carrier wave in a propagated signal, where the computer-readable program code can be loaded onto an instruction execution system, apparatus, or device. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF, or any suitable combination thereof.
[0199] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0200] Another embodiment of the present invention provides a computer program product, which, when executed, is used to perform the above-described personalized live streaming method.
[0201] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments of the present invention.
[0202] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in this invention is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms for implementing the invention.
[0203] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0204] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions in the present invention.
Claims
1. A personalized live streaming method, characterized in that, include: Real-time acquisition of on-site image data based on camera equipment; A virtual agent mapping table is generated based on the on-site image data; Receive a user's virtual agent selection request; wherein, the virtual agent selection request includes the target virtual agent selected by the user and user terminal information; Based on the identifier of the target virtual seat, the target clipping information is retrieved from the virtual seat mapping table; The target cropping information is sent to the user terminal; wherein, the user terminal performs real-time image processing on the current live frame based on the target cropping information to generate the terminal display screen.
2. The personalized live streaming method according to claim 1, characterized in that, The step of generating a virtual agent mapping table based on the on-site image data includes: The valid seats in the on-site image data are determined based on a preset seat structure template; For each valid seat, record the coordinates of the center point of the valid seat and assign a unique number to the valid seat; Based on the center point coordinates of the valid seats, the coordinates of the valid seats are converted into virtual coordinates in the virtual interface; The structured data of the valid seat is generated based on the unique number of the valid seat, the coordinates of the valid seat, the virtual coordinates of the valid seat, the camera number of the area to which the valid seat belongs, and the image cropping area. A virtual agent mapping table is generated based on the structured data of all the valid agents.
3. The personalized live streaming method according to claim 2, characterized in that, After determining the valid seats in the on-site image data based on the preset seat structure template, the method further includes: Human figure recognition is performed on the area where the valid seats are located to obtain a human figure recognition result; wherein, the human figure recognition result indicates whether there is a person in the area where the valid seats are located. Based on the human figure recognition results, personnel information is marked for the valid seats in the virtual interface.
4. The personalized live streaming method according to claim 1, characterized in that, After generating the virtual agent mapping table based on the on-site image data, the method further includes: Perform a consistency check on the virtual agent mapping table to obtain the consistency check result; For each abnormal region in the consistency detection result, the virtual coordinates in the abnormal region are corrected.
5. The personalized live streaming method according to claim 1, characterized in that, The position and angle of each camera device are determined based on the site layout plan and the viewing angle parameters of the camera device; wherein, the site layout plan includes seating distribution information and stage position; the viewing angle parameters of the camera device include focal length, horizontal viewing angle and vertical viewing angle.
6. The personalized live streaming method according to claim 2, characterized in that, Also includes: If the area where the seat is located is covered by multiple camera devices at the same time, the images captured by the multiple camera devices are merged to obtain a merged image.
7. The personalized live streaming method according to claim 1, characterized in that, Also includes: The priority score of the camera device is determined based on the number of times the target virtual seat is selected, the number of seats covered by the camera device, and the set of image cropping areas covered by the camera device. Image processing resources and transmission bandwidth are then dynamically allocated based on the priority score of the camera device.
8. A personalized live streaming device, characterized in that, include: The acquisition unit is used to acquire real-time on-site image data based on camera equipment; The generation unit is used to generate a virtual agent mapping table based on the on-site image data; A receiving unit is configured to receive a user's virtual agent selection request; wherein the virtual agent selection request includes the target virtual agent selected by the user and user terminal information; The retrieval unit is used to retrieve target clipping information from the virtual seat mapping table based on the identifier of the target virtual seat; The sending unit is used to send the target cropping information to the user terminal; wherein the user terminal performs real-time image processing on the current live frame according to the target cropping information to generate the terminal display screen.
9. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the personalized live streaming method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, It stores a computer program, wherein the computer program, when executed by a processor, implements the personalized live streaming method as described in any one of claims 1 to 7.