Media item management method and apparatus, device, and medium
By obtaining the target location in the live broadcast application and determining matching media items from multiple acquisition devices, the problem of viewers being unable to choose the viewing location is solved, and flexible live video viewing and better visual effects are achieved.
Patent Information
- Application Number
- PCT/CN2025/084840
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-02
AI Technical Summary
In existing live broadcast applications, viewers cannot select a viewing position other than a fixed position, resulting in unsatisfactory panoramic image visual effects and image deformation problems.
By obtaining a target position, determining media items matching the position from multiple acquisition devices based on the position, generating a target media item for the audience to watch at any viewpoint position, using virtual space interaction to select the viewpoint position, and generating the target media item by weighted summation of multiple media items.
It enables viewers to watch flexibly from any viewpoint, provides more accurate and rich visual information, reduces transmission bandwidth and delay, and reduces server load.
Smart Images

Figure CN2025084840_02102025_PF_FP_ABST
Abstract
Description
Method, apparatus, device, and medium for managing media items
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, apparatus and media for managing media items” and application number 2024103847032, filed on March 29, 2024, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Exemplary implementations of the present disclosure relate generally to media management, and more particularly to methods, apparatuses, devices, and computer-readable storage media for managing media items in a live broadcast application. Background Art
[0003] With the advancement of computer and network technologies, a variety of live streaming applications have been developed. These applications can provide live streaming services to a large audience in real time. Furthermore, multi-location live streaming technology solutions have been proposed. For example, multiple acquisition devices can be deployed in a live streaming environment to provide viewers with live video from multiple viewpoints. However, because the deployment locations of multiple acquisition devices are fixed, viewers cannot choose viewing positions outside of these fixed locations. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for managing media items is provided. In this method, a target location for viewing a live broadcast object is determined. Based on the target location, a set of media items is obtained, each of which is from a set of capture devices associated with the live broadcast object. Based on the set of media items, a target media item matching the target location is determined.
[0005] In a second aspect of the present disclosure, a device for managing media items is provided. The device includes: a location acquisition module configured to acquire a target location for viewing a live broadcast object; a media acquisition module configured to acquire a set of media items based on the target location, the set of media items being from a set of acquisition devices associated with the live broadcast object; and a determination module configured to determine, based on the set of media items, a target media item that matches the target location.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1A shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0011] FIG1B shows a block diagram of deployment locations of multiple acquisition devices according to an exemplary implementation of the present disclosure;
[0012] FIG2 illustrates a block diagram for managing media items according to some implementations of the present disclosure;
[0013] FIG3 illustrates a block diagram for determining a target location according to some implementations of the present disclosure;
[0014] FIG4 shows a block diagram for determining a group of acquisition devices according to some implementations of the present disclosure;
[0015] FIG5 shows a block diagram for determining a group of acquisition devices according to some implementations of the present disclosure;
[0016] 6 illustrates a block diagram for determining pixels in a target media item according to some implementations of the present disclosure;
[0017] FIG7 illustrates a block diagram for scaling a target media item according to some implementations of the present disclosure;
[0018] FIG8 illustrates a flowchart of a method for managing media items according to some implementations of the present disclosure;
[0019] FIG9 illustrates a block diagram of an apparatus for managing media items according to some implementations of the present disclosure; and
[0020] FIG10 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0021] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0022] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
[0029] Sample Environment
[0030] A variety of live streaming applications have been developed. Live streaming services can be provided to a large number of viewers in real time via live streaming applications. Referring to FIG. 1A , an application environment according to an example implementation of the present disclosure is described. FIG. 1A shows a block diagram 100A of an application environment according to an example implementation of the present disclosure. As shown in FIG. 1A , a first client device 112 of a first user 110 of the live streaming application can access a server device 130 via a network 140, and a second user 120 can access the server device 130 using a second client device 122.
[0031] Here, the first user 110 can be, for example, a viewer in a live broadcast application, and the second user 120 can be, for example, a host and / or guest in the live broadcast application. According to an example implementation of the present disclosure, the second user 120 can display a live broadcast object (e.g., a certain product) and introduce various aspects of the product. Alternatively and / or additionally, the second user 120 can serve as a live broadcast object and broadcast a talent show (e.g., singing, dancing, etc.). For ease of description, the live broadcast process will be described below using the second user as the live broadcast object as an example.
[0032] 1A , a plurality of collection devices 124, 126, ..., 128 may be deployed at various locations around the second user 120. The second client device 122 may send a plurality of media items from the plurality of collection devices to the server device 130 via the network 140, so that the server device 130 provides the respective media items to the first client devices 112 of one or more first users 110.
[0033] Refer to Figure 1B for more details about deploying acquisition devices, which shows a block diagram 100B of the deployment positions of multiple acquisition devices according to an exemplary implementation of the present disclosure. As shown in Figure 1B, the position of the second user 120 can be used to establish an xyz coordinate system as the origin O, and the direction of the xyz axis is shown in Figure 1B. Multiple acquisition devices can be deployed at points P1, P2, P3, P4, and P5, respectively, and the direction of each acquisition device is toward the origin. For example, the direction of the acquisition device at P1 is the direction shown in P1->O, the direction of the acquisition device at P2 is the direction shown in P2->O, and so on. In this way, multiple acquisition devices can provide live video from multiple angles to the audience. At this time, the first user 110 can choose to watch media items from a certain (or some) acquisition device. However, since the deployment positions of multiple acquisition devices are fixed, the audience cannot choose a viewing position outside these fixed positions.
[0034] Currently, panorama-based roaming technology solutions have been proposed. These can provide a panoramic image generated based on multiple media items to the first user 110 and allow the user to select different viewpoints. However, the visual effect of panoramic images is not ideal, and the image presented at certain locations may be severely distorted. Therefore, it is desirable to provide users with videos from any viewpoint, alleviate the image distortion problem in such videos, and provide more accurate and realistic videos.
[0035] Summary of Managing Media Items
[0036] To at least partially address the deficiencies in the prior art, a method for managing media items is provided according to an exemplary implementation of the present disclosure. An overview of an exemplary implementation of the present disclosure is described with reference to FIG2 , which illustrates a block diagram 200 for managing media items according to some implementations of the present disclosure. As shown in FIG2 , a target location 210 for viewing a live object can be obtained. This target location 210 may also be referred to as a user's viewing location or viewpoint location.
[0037] According to an example implementation of the present disclosure, media items may include at least any one of images and / or videos. In the following, the process of media item management will be described using only videos as an example. In this way, it is possible to support the provision of richer media data to users. Based on the target location, a group of media items 230 can be obtained, each of which comes from a group of acquisition devices associated with the live broadcast object. A group of acquisition devices may be at least a portion of a plurality of acquisition devices. A plurality of acquisition devices (e.g., acquisition devices 124, 126, ..., and 128) are associated with the target object and are deployed within a predetermined spatial range of the second user (i.e., the target object). Here, the predetermined spatial range may represent a pre-specified spatial range around the live broadcast object, for example, the room where the live broadcast object is located, or a spatial range within 3 meters (or other values) around the live broadcast object.
[0038] For example, a group of capture devices may be determined based on the distances between the locations of the multiple capture devices and the target location 210, or based on the distances between a straight line including the directions of the multiple capture devices and the target location 210. A target media item 240 matching the target location 210 may be determined using the group of media items. Furthermore, the target media item 240 may be provided to the first client device 112 so that the first user 110 can view the live video at any viewpoint.
[0039] Utilizing the exemplary implementations of the present disclosure, users can be provided with more flexible control methods, allowing them to select any desired viewpoint. Furthermore, media items from relevant capture devices among multiple capture devices can be used to generate a video at that viewpoint. Using the exemplary implementations of the present disclosure, the determined media items are generated using a set of media items captured in real time, and thus can include more accurate and rich visual information. In this way, users can view the live broadcast from any desired viewpoint during the live broadcast.
[0040] Detailed procedures for managing media items
[0041] Having described an overview of an example implementation of the present disclosure, further details regarding managing media items are described below with reference to the accompanying drawings. According to an example implementation of the present disclosure, a target location may indicate the location of a second user, from which a first user is to view a live broadcast application (i.e., the first user's viewpoint location). For example, a selection page may be provided on the first client device 112 so that the first user 110 can select a desired viewpoint location via the selection page. In this case, the target location may be determined based on the first user's input.
[0042] During the process of acquiring the target location, a virtual space may be presented to the first user 110 at the first client device 112. The virtual space may indicate multiple virtual device locations corresponding to multiple device locations of the multiple acquisition devices. Furthermore, the first user may interact with the virtual space to determine the target location corresponding to the interaction operation based on the interaction operation.
[0043] For more details on determining the target position, see FIG3 , which shows a block diagram 300 for determining the target position according to some implementations of the present disclosure. As shown in FIG3 , the virtual space can present the positions of multiple acquisition devices and the position of the second user 120 in three-dimensional coordinates. For example, the second user 120 can be located at the coordinate origin O, and the multiple acquisition devices can be located at multiple points P1, P2, P3, P4, and P5 at the x, y, and z coordinate axes, respectively. Further, the user can perform an interactive operation 320 to select a target position 310 (for example, represented by coordinates (x0, y0, z0)) at which the live video is desired to be viewed. For example, the first user can determine the target position 310 by single-clicking, double-clicking, dragging, or the like. Using the example implementation of the present disclosure, a selection page can be provided to the user in a visual manner so that the user can select the desired viewpoint position in a more convenient and accurate manner.
[0044] It should be understood that although FIG3 schematically illustrates an example in which multiple acquisition devices are deployed at respective coordinate axes, alternatively and / or additionally, the acquisition devices can be deployed at locations other than the coordinate axes. For example, acquisition devices can be deployed at any location on the sphere shown in FIG3 along a direction toward the origin. In this way, media items from more angles can be provided, thereby improving the accuracy of generating target media items. Alternatively and / or additionally, the distances between the multiple acquisition devices and the origin O can be different. In this way, long-distance images and / or close-up images can be presented according to user needs.
[0045] According to an exemplary implementation of the present disclosure, a second client device can transmit multiple media items from multiple capture devices to a server device. For example, the multiple media items can be transmitted individually to the server device. Alternatively and / or additionally, encoding operations can be performed on the multiple media items so that the multiple media items are transmitted as a single data stream. In this way, the bandwidth involved in the transmission process can be reduced, transmission delay can be reduced, and transmission efficiency can be improved.
[0046] According to an example implementation of the present disclosure, a server device of a live broadcast application can determine which media items to provide to a first client device based on a target location. In this way, the first client device can directly receive the required media items from the server device, thereby reducing the workload of the first client device.
[0047] According to an example implementation of the present disclosure, a server device can obtain a set of directions for a group of collection devices. Here, for the first collection device in the group, the first direction of the first collection device is from the first position of the first collection device to the position of the second user. Specifically, as shown in Figure 3, the directions of each collection device at points P1 to P5 all point to the origin O. Furthermore, the space defined by the target location and the directions of each collection device can be compared to determine a group of collection devices.
[0048] In the context of the present disclosure, a vector passing through the position of the acquisition device can be used to represent the direction of the acquisition device. There may be various spatial relationships between the target position and the position of the acquisition device: for example, the target position may be located in a one-dimensional space (i.e., a straight line) defined by the position of the acquisition device and the position of the second user, the target position may be located in a two-dimensional space (i.e., a plane) defined by the positions of two acquisition devices and the position of the second user, or the target position may be located in a three-dimensional space (i.e., a volume space) defined by the positions of three acquisition devices and the position of the second user.
[0049] 4 , which illustrates a block diagram 400 for determining a group of collection devices according to some implementations of the present disclosure. For ease of description, assume that three collection devices are deployed at positions P1, P2, and P5, respectively. The collection device at P1 (e.g., referred to as the first collection device) is oriented in the -x direction, the collection device at P2 (e.g., referred to as the second collection device) is oriented in the -y direction, and the collection device at P5 (e.g., referred to as the third collection device) is oriented in the -z direction.
[0050] According to an exemplary implementation of the present disclosure, assuming that the target position is at position 410, the target position is located in a one-dimensional space (i.e., the x-axis) defined by a first position of a first acquisition device (i.e., point P1) and a position of a second user (i.e., origin O). In this case, a set of media items may include media items acquired by the first acquisition device.
[0051] It should be understood that FIG4 is merely illustrative, and the target position may be located elsewhere. Assuming the target position is located on the y-axis, a set of media items may include media items from the capture device at point P2, and so on. In this way, by comparing the positional relationship between the target position and the straight line defined by the position of the capture device and the position of the second user, the data required to generate the target media item can be determined in a simple and accurate manner.
[0052] According to an exemplary implementation of the present disclosure, assuming that the target position is at position 420, the target position is located in a two-dimensional space (i.e., the xoy plane) defined by a first position of a first acquisition device (i.e., point P1), a second position of a second acquisition device (i.e., point P2), and a position of a second user (i.e., origin O). In this case, a set of media items may include media items collected by the first acquisition device and media items collected by the second acquisition device.
[0053] It should be understood that FIG4 is merely illustrative, and the target location can be located elsewhere. Assuming the target location is located in the xoz plane, a set of media items can include media items from a capture device at point P1 and media items from a capture device at point P5, and so on. In this way, by comparing the positional relationship between the target location and the plane defined by the positions of the two capture devices and the position of the second user, the data required to generate the target media item can be determined in a simple and accurate manner.
[0054] More examples of target locations outside of the coordinate plane are described with reference to FIG5 , which illustrates a block diagram 500 for determining a group of capture devices according to some implementations of the present disclosure. As shown in FIG5 , assume that the target location is at position 510. At this time, position 510 is located in a three-dimensional space defined by a first position of a first capture device (i.e., point P1), a second position of a second capture device (i.e., point P2), a third position of a third capture device (i.e., point P5), and a position of a second user (i.e., origin O) (i.e., the first quadrant defined by origin O, +x-axis, +y-axis, and +z-axis). At this time, a group of media items may include media items captured by the first capture device, media items captured by the second capture device, and media items captured by the third capture device.
[0055] It should be understood that FIG5 is merely schematic, and the target position may be located elsewhere. Assuming the target position is located in a space defined by the origin O, the -x axis, the +y axis, and the +z axis, a set of media items may include media items from capture devices at points P2, P3, and P5, and so on. In this way, by comparing the positional relationship of the target position and the plane defined by the positions of the three capture devices and the position of the second user, the data required to generate the target media item can be determined in a simple and accurate manner.
[0056] According to an exemplary implementation of the present disclosure, a server device may generate target media items based on a target location and a set of media items and transmit the generated media items to client devices of respective viewers. It should be understood that during a live broadcast, there may be a large number of viewers, and generating target media items at the server device may result in an excessive workload on the server device.
[0057] According to an example implementation of the present disclosure, a server device may transmit a determined set of media items to a first client device. For example, each media item may be transmitted separately to the server device. Alternatively and / or additionally, each media item may be encoded so as to be transmitted as a single data stream. In this way, the bandwidth involved in the transmission process may be reduced, transmission delay may be reduced, and transmission efficiency may be improved. The client device may receive a set of media items from the server device and generate target media items using the set of media items. In this way, the workload of the server device may be reduced, and the server device may be able to focus on other computational processing processes of the live broadcast process.
[0058] According to an exemplary implementation of the present disclosure, a first client device can determine a set of weights based on the distance between a target location and a set of directions of a set of capture devices. Subsequently, target media items can be determined based on the set of weights and the set of media items. Returning to FIG. 4 , assuming the target location is at position 410 in FIG. 4 , the server device only transmits media items from the capture device at point P1 to the first client device. At this point, the distance between the target location and the x-axis is zero, so the media items from the capture device at point P1 can be directly used as target media items.
[0059] Assuming the target location is at location 420 in Figure 4 , the server device can transmit two media items from the capture devices at points P1 and P2 to the first client device. In this case, the corresponding weights can be determined based on the ratio between location 420 and the x-axis and y-axis. Assuming location 420 is at the angle bisector of the x-axis and y-axis (i.e., 45 degrees), the contributions from the two media items are equal, and the specific content of the target media item can be determined based on a weighted summation.
[0060] Assuming the target position is at position 510 in FIG5 , the distance between position 510 and the directions of each acquisition device (i.e., +x axis, +y axis, and +z axis) can be determined to determine the corresponding weight. Specifically, assuming the coordinates of position 510 are (x0, y0, z0), the distance 416 between position 510 and the +x axis is expressed as The distance 412 between the position 510 and the +y axis is represented as The distance 414 between the position 510 and the +z axis is represented as In this way, determining the contribution of each media item to the final target media item can be converted into a mathematical operation process, thereby determining the target media item in a simpler and more accurate manner.
[0061] According to an example implementation of the present disclosure, the weights of the various media items can be determined based on the proportional relationship between the aforementioned distances. Specifically, in the process of generating a target media item, for a target pixel in the target media item, a group of pixels corresponding to the target pixel in a group of media items is determined. Furthermore, the color information of the target pixel can be determined based on a set of weights and the color information of a group of pixels. For more information, see FIG6 , which shows a block diagram 600 for determining pixels in a target media item according to some implementations of the present disclosure.
[0062] As shown in FIG6 , assuming that media items from three capture devices are represented as media items 610 (media item from the capture device at point P1), 620 (media item from the capture device at point P2), and 630 (media item from the capture device at point P5), in the process of determining the color of pixel 652 (e.g., pixel coordinates (i, j)) in target media item 650, a weighted sum 640 can be performed on pixel 612 in media item 610, pixel 622 in media item 620, and pixel 632 in media item 630. It should be understood that the pixel coordinates of pixels 612, 622, and 624 are all (i, j), and each pixel position can be traversed to determine the corresponding color. For example, the corresponding color can be determined based on the following formula 1.
[0063] In the above formula, color represents the color data of pixel 652, color x Indicates the color data of pixel 612, color y Indicates the color data of pixel 622, and color z represents the color data of pixel 632. In this way, the color data of each pixel in the target media item can be determined based on simple mathematical operations. It should be understood that the above formula is merely exemplary, and alternatively and / or additionally, a media item with a viewpoint at a desired position and / or orientation can be generated based on multiple media items from capture devices deployed at different positions and / or orientations based on various technical solutions currently known and / or to be developed in the future.
[0064] According to an example implementation of the present disclosure, when the media item is an image, a group of images can be used to generate a final image based on the method described above. When the media item is a video, a group of images with the same timestamp can be used to generate an image associated with that timestamp. For example, a group of images at time point t can be used to generate an image at time point t, and a group of images at time point t+1 can be used to generate an image at time point t+1. The images can then be presented in chronological order.
[0065] According to an example implementation of the present disclosure, in the process of generating a target media item, the target media item can be scaled based on the distance between the target position and the position of the second user. See Figure 7 for more details on the scaling operation, which shows a block diagram 700 for scaling a target media item according to some implementations of the present disclosure. As shown in Figure 7, assuming that the target position is located at position 710, that is, the distance between the target position and the origin O is less than the distance between the acquisition device at point P2 and the origin O, at this time, the image content in the target media item can be amplified. For another example, assuming that the target position is located at position 720, that is, the distance between the target position and the origin O is less than the distance between the acquisition device at point P2 and the origin O, at this time, the image content in the target media item can be reduced. In this way, users can be supported to select viewpoint positions in a more flexible manner, thereby obtaining more information about live broadcasts.
[0066] It should be understood that the above description is merely an example of a process where the target position is located in the first quadrant of the three-dimensional space. Alternatively and / or additionally, the target position may be located in other quadrants of the three-dimensional space. The corresponding weights may be determined based on the spatial geometric relationship between the positions of the various points, which will not be further described below.
[0067] It should be understood that the above description only appears to illustrate a situation where three acquisition devices are located on the +x, +y, and +z axes, respectively. Alternatively and / or additionally, the acquisition devices can be located at other locations in three-dimensional space. In this case, a corresponding coordinate system can be established with the second user's position as the origin, and the specific expression of the direction and position of each acquisition device can be determined based on the spatial positional relationship between the position of each acquisition device and the origin. Furthermore, according to the principles described above, a group of acquisition devices can be determined from multiple acquisition devices, and then the final target media item can be determined based on the media items from the group of acquisition devices.
[0068] According to an example implementation of the present disclosure, the first user can be a viewer in a live broadcast application, and the live broadcast object can be a second user in the live broadcast application, such as at least one of the host or guest. In this case, ordinary viewers can select any viewpoint to watch the live broadcast. Using this example implementation of the present disclosure, the generated video is generated using a set of media items collected in real time, and thus can include more accurate and rich visual information. In this way, users can watch the live broadcast from any desired viewpoint during the live broadcast.
[0069] Example Process
[0070] FIG8 illustrates a flow chart of a method 800 for managing media items according to some implementations of the present disclosure. At block 810, a target location for viewing a live broadcast object is obtained. At block 820, a set of media items is obtained based on the target location, the set of media items being from a set of capture devices associated with the live broadcast object. At block 830, a target media item matching the target location is determined based on the set of media items.
[0071] According to an example implementation of the present disclosure, obtaining a target position includes: presenting a virtual space, the virtual space indicating multiple virtual device positions corresponding to multiple device positions of multiple acquisition devices, respectively, the multiple acquisition devices are associated with a live broadcast object, and a group of acquisition devices is at least a part of the multiple acquisition devices; and in response to an interactive operation of a first user in the virtual space, determining a target position corresponding to the interactive operation.
[0072] According to an example implementation of the present disclosure, determining a target media item includes: obtaining a set of directions of a set of acquisition devices, wherein, for a first acquisition device in the set of acquisition devices, a first direction of the first acquisition device points from a first position of the first acquisition device to a position of a live broadcast object; determining a set of weights based on a distance between the target position and a set of straight lines defined by the set of directions; and generating a target media item based on the set of weights and a set of media items.
[0073] According to an example implementation of the present disclosure, a media item includes at least any one of an image and a video, and generating a target media item includes: determining, for a target pixel in the target media item, a set of pixels in a set of media items that correspond to the target pixel; and determining color information of the target pixel based on a set of weights and color information of a set of pixels.
[0074] According to an example implementation of the present disclosure, generating the target media item further includes scaling the target media item based on a distance between the target position and a position of the live object.
[0075] According to an example implementation of the present disclosure, a target location is located within a space defined by a set of directions of a set of acquisition devices.
[0076] According to an exemplary implementation of the present disclosure, the target position is located in a one-dimensional space defined by the first position of the first acquisition device and the position of the live broadcast object.
[0077] According to an example implementation of the present disclosure, a group of collection devices further includes a second collection device, and the target position is located in a two-dimensional space defined by a first position of the first collection device, a second position of the second collection device, and a position of the live object.
[0078] According to an example implementation of the present disclosure, a group of acquisition devices further includes a second acquisition device and a third acquisition device, and the target position is located in a three-dimensional space defined by a first position of the first acquisition device, a second position of the second acquisition device, a third position of the third acquisition device, and a position of the live broadcast object.
[0079] According to an example implementation of the present disclosure, the method is performed at a first client device of a first user, and a set of media items is determined by a server device of a live broadcast application based on a target location.
[0080] According to an example implementation of the present disclosure, obtaining a target location includes: determining the target location based on the input of a first user in a live broadcast application, the first user is a viewer in the live broadcast application, the live broadcast object is a second user in the live broadcast application, and the second user is at least any one of a host or a guest, and the method further includes: presenting a target media item at a first client device.
[0081] Example devices and equipment
[0082] Figure 9 shows a block diagram of an apparatus 900 for managing media items according to some implementations of the present disclosure. The apparatus includes: a location acquisition module 910 configured to acquire a target location for viewing a live broadcast object; a media acquisition module 920 configured to acquire a set of media items based on the target location, the set of media items being from a set of acquisition devices associated with the live broadcast object; and a determination module 930 configured to determine a target media item that matches the target location based on the set of media items.
[0083] According to an example implementation of the present disclosure, a position acquisition module includes: a presentation module configured to present a virtual space, the virtual space indicating multiple virtual device positions corresponding to multiple device positions of multiple acquisition devices, respectively, the multiple acquisition devices are associated with a live broadcast object, and a group of acquisition devices is at least a part of the multiple acquisition devices; and a position determination module configured to determine, in response to an interactive operation of a first user in the virtual space, a target position corresponding to the interactive operation.
[0084] According to an example implementation of the present disclosure, a determination module includes: a direction acquisition module, configured to acquire a set of directions of a group of acquisition devices, wherein, for a first acquisition device in a group of acquisition devices, a first direction of the first acquisition device points from a first position of the first acquisition device to a position of a live broadcast object; a weight determination module, configured to determine a set of weights based on a distance between a target position and a set of straight lines defined by a set of directions; and a generation module, configured to generate target media items based on a set of weights and a set of media items.
[0085] According to an example implementation of the present disclosure, the media item includes at least any one of an image and a video, and the generation module includes: a pixel determination module configured to determine, for a target pixel in a target media item, a group of pixels in a group of media items corresponding to the target pixel; and a color determination module configured to determine color information of the target pixel based on a set of weights and color information of a group of pixels.
[0086] According to an exemplary implementation of the present disclosure, the generating module further includes: a scaling module configured to scale the target media item based on a distance between the target position and a position of the live object.
[0087] According to an example implementation of the present disclosure, a target location is located within a space defined by a set of directions of a set of acquisition devices.
[0088] According to an exemplary implementation of the present disclosure, the target position is located in a one-dimensional space defined by the first position of the first acquisition device and the position of the live broadcast object.
[0089] According to an example implementation of the present disclosure, a group of collection devices further includes a second collection device, and the target position is located in a two-dimensional space defined by a first position of the first collection device, a second position of the second collection device, and a position of the live object.
[0090] According to an example implementation of the present disclosure, a group of acquisition devices further includes a second acquisition device and a third acquisition device, and the target position is located in a three-dimensional space defined by a first position of the first acquisition device, a second position of the second acquisition device, a third position of the third acquisition device, and a position of the live broadcast object.
[0091] According to an example implementation of the present disclosure, the apparatus is implemented at a first client device of a first user, and a set of media items is determined by a server device of a live broadcast application based on a target location.
[0092] According to an example implementation of the present disclosure, the location acquisition module is further configured to: determine the target location based on the input of a first user in the live broadcast application, the first user is a viewer in the live broadcast application, the live broadcast object is a second user in the live broadcast application, and the second user is at least any one of a host or a guest, and the apparatus further includes: a presentation module configured to present the target media item at the first client device.
[0093] FIG10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in FIG10 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 1000 shown in FIG10 can be used to implement the methods described above.
[0094] As shown in FIG10 , computing device 1000 is in the form of a general-purpose computing device. Components of computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit 1010 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1000.
[0095] The computing device 1000 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1020 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1000.
[0096] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0097] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0098] Input device 1050 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1060 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1000 may also communicate with one or more external devices (not shown) via communication unit 1040 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1000, or with any device that allows computing device 1000 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0099] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0100] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0101] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0102] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0103] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0104] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for managing media items, comprising: Get the target location for viewing the live object; Based on the target location, obtaining a group of media items, wherein the group of media items are respectively from a group of acquisition devices associated with the live broadcast object; as well as A target media item matching the target location is determined based on the set of media items.
2. The method according to claim 1, wherein obtaining the target position comprises: presenting a virtual space, the virtual space indicating a plurality of virtual device positions corresponding to a plurality of device positions of a plurality of acquisition devices, the plurality of acquisition devices being associated with the live broadcast object, and the group of acquisition devices being at least a portion of the plurality of acquisition devices; as well as In response to an interaction operation of the first user in the virtual space, the target position corresponding to the interaction operation is determined.
3. The method of claim 1 , wherein determining the target media item comprises: Acquire a set of directions of the group of acquisition devices, where for a first acquisition device in the group of acquisition devices, a first direction of the first acquisition device points from a first position of the first acquisition device to a position of the live broadcast object; determining a set of weights based on distances between the target location and a set of lines defined by the set of directions; as well as The target media item is generated based on the set of weights and the set of media items.
4. The method according to claim 3, wherein the media item comprises at least any one of an image and a video, and generating the target media item comprises: For a target pixel in the target media item, determining a group of pixels in the group of media items that correspond to the target pixel; as well as Color information of the target pixel is determined based on the set of weights and color information of the set of pixels.
5. The method of claim 3, wherein generating the target media item further comprises: The target media item is scaled based on a distance between the target location and a location of the live object. The method of claim 3 , wherein the target location is located within a space defined by the set of directions of the set of acquisition devices.
7. The method according to claim 6, wherein the target position is located in a one-dimensional space defined by the first position of the first acquisition device and the position of the live broadcast object.
8. The method according to claim 6, wherein the group of acquisition devices further includes a second acquisition device, and the target position is located in a two-dimensional space defined by a first position of the first acquisition device, a second position of the second acquisition device, and a position of the live broadcast object.
9. The method according to claim 6, wherein the group of acquisition devices further includes a second acquisition device and a third acquisition device, and the target position is located in a three-dimensional space defined by the first position of the first acquisition device, the second position of the second acquisition device, the third position of the third acquisition device, and the position of the live broadcast object.
10. The method of claim 1, wherein the method is performed at a first client device of the first user, and the set of media items is determined by a server device of the live broadcast application based on the target location.
11. The method according to claim 10, wherein obtaining the target position comprises: The target location is determined based on the input of a first user in the live broadcast application, the first user is a viewer in the live broadcast application, the live broadcast object is a second user in the live broadcast application, and the second user is at least any one of a host or a guest, and the method further includes: presenting the target media item at the first client device.
12. An apparatus for managing media items, comprising: A location acquisition module configured to acquire a target location for viewing a live broadcast object; a media acquisition module configured to acquire a group of media items based on the target location, the group of media items respectively coming from a group of acquisition devices associated with the live broadcast object; as well as The determining module is configured to determine a target media item matching the target position based on the group of media items.
13. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.
14. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Parallax generation method, generation cell and three-dimensional video generation method and device
CN101321299A
Image reconstruction method and device, computer readable storage medium and processor
CN114092315A
Image processing apparatus, image processing method, and image processing program
JP2020043467A
Three-dimensional content distribution system, three-dimensional content distribution method, and computer program
JP2020127211A