Information processing device, information processing method, and program

The information processing apparatus addresses the challenge of real-time performer identification and dynamic posture prediction by integrating person identification and PTZ camera control to create content that accurately combines live performer actions with CG characters, improving the viewing experience.

WO2026070293A1PCT designated stage Publication Date: 2026-04-02SONY GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing technologies struggle to appropriately identify performers in live performances and incorporate additional elements, such as computer-generated characters, due to challenges in real-time person identification and dynamic posture prediction during vigorous movements.

Method used

An information processing apparatus and method that includes a calculation unit to analyze captured images, predict the posture and position of performers, generate control signals for PTZ cameras to adjust imaging direction and zoom, and integrate person identification information with motion data to generate content images featuring CG characters.

Benefits of technology

Enables accurate and timely imaging of performers, allowing for the creation of content that seamlessly integrates real-time performer actions with CG characters, enhancing the viewing experience by ensuring clear and appropriate capture of performers' faces and bodies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025031398_02042026_PF_FP_ABST
    Figure JP2025031398_02042026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises a calculation unit that performs posture prediction processing of analyzing a captured image including a performer as a subject and predicting a position or a posture of the performer after the current time, and control signal generation processing of generating a control signal for changing an imaging direction or a zoom magnification of an imaging device by using the face of the performer as a target according to a prediction result of the posture prediction processing.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, Program

[0001] The present technology relates to an information processing apparatus, an information processing method, and a program, and particularly relates to a technology related to imaging of a person.

[0002] In recent years, services for distributing various contents through networks such as the Internet have been widely developed. For example, Patent Document 1 discloses a technology for improving a user's viewing experience in content distribution.

[0003] Japanese Patent Application Laid-Open No. 2021-192470

[0004] By the way, in the current situation where various entertainment contents such as music concerts, plays, and movies are widely distributed, there is a desire to appropriately identify performers and incorporate additional elements accordingly.

[0005] Therefore, in the present disclosure, a technology that enables appropriate imaging of performers in music concerts and the like and can be used for, for example, person identification is proposed.

[0006] The information processing apparatus according to the present technology includes an arithmetic unit that performs a posture prediction process of analyzing a captured image with a performer as a subject and predicting the position or posture of the performer after the current time, and a control signal generation process of generating a control signal for changing the imaging direction or zoom magnification of an imaging device targeting the face of the performer according to the prediction result of the posture prediction process. It estimates the position or posture of the performer after the current time, and at the time when the performer reaches that position or posture, the imaging direction and zoom magnification of the imaging device are changed so that the performer can be appropriately imaged.

[0007] Explanatory diagram of the content providing system according to an embodiment of the present technology. Explanatory diagram of a live venue when applying the content providing system according to the embodiment. Explanatory diagram of motion data in the embodiment. Explanatory diagram of the content screen provided in the embodiment. Explanatory diagram of the association between a performer and a character in the embodiment. Block diagram of the information processing apparatus according to the embodiment. Explanatory diagram of the processing of the arithmetic unit according to the embodiment. Explanatory diagram of the information associated with the tracking identification information according to the embodiment. Flowchart of the priority setting by the arithmetic unit according to the embodiment.

[0008] The embodiments will be described below in the following order: <1. System Configuration> <2. Server Device Configuration> <3. Processing Examples> <4. Summary and Modifications>

[0009] In this disclosure, "image" refers to both video and still images. Furthermore, "image" refers to the image actually displayed on the screen, but "image" in the signal processing process and transmission path leading up to the display on the screen refers to image data. In the service of the embodiment, content is mainly provided as images, including still images, videos, and pseudo-videos made up of a small number of still images. Some of these images include sound, while others do not.

[0010] Furthermore, "performer" refers to someone who performs some kind of act in front of an audience, and in this embodiment, one example is a singer or dancer at a music live performance.

[0011] <1. System Configuration> Figure 1 shows an overview of the content provision system 1 of this embodiment. In this embodiment, an example is a service that provides users with images of computer graphics (CG) characters of a live music performance by a group of real people. For example, each singer in an actual singing group is associated with a CG character on a one-to-one basis. The image content then shows the CG characters corresponding to each singer performing songs and dances based on live images of the singing group. However, the content provision system 1 of this embodiment can be applied to performance venues other than live music performances, and its application to live music performances is just one example.

[0012] Figure 1 shows the hardware elements that constitute the content provision system 1, including a server device 2, a database unit 3 (hereinafter referred to as "DB"), a local controller 4, a player device 5, an imaging device 6 (hereinafter referred to as "camera 6"), and a pan-tilt-zoom variable imaging device 7 (hereinafter referred to as "PTZ camera 7"). Note that these are just examples of the hardware elements that constitute the content provision system 1, and some of the components may be omitted, or hardware elements other than those shown may be added.

[0013] In the content provision system 1 shown in Figure 1, the server device 2, local controller 4, and player device 5 are able to communicate with each other via, for example, a wired or wireless public network or a dedicated network. Alternatively, these devices may communicate via wired or wireless communication between them instead of network communication.

[0014] This content provision system 1 is a system that provides a service (hereinafter referred to as the "performance collaboration service") that offers CG image content based on the performance of performers such as singers to users who have a player device 5. Users can view the content images on the player device 5.

[0015] The multiple cameras 6 in the content provision system 1 are configured as digital camera devices having image sensors such as CCD (Charge Coupled Devices) sensors or CMOS (Complementary Metal-Oxide-Semiconductor) sensors, and acquire captured images as digital data. For example, each camera 6 acquires captured images as video.

[0016] This camera 6 primarily captures images of the stage where performers are located from various positions with a relatively wide angle, for example, in a live music venue. For example, Figure 2 shows a live music venue 15. The live venue 15 is broadly divided into a stage 16 where performers 19 such as singers stand, and a floor 17 where the audience is located. In such a live venue 15, multiple cameras 6 are arranged so that they can capture images of the stage 16 from various positions.

[0017] For example, the content provision system 1 extracts motion data of each performer 19, who are the subjects of the video, from the images captured by the camera 6. Motion data is skeletal capture data that allows the system to recognize the skeletal position of the performer 19, including information such as the central position coordinates of the performer 19's body and the position coordinates of each joint in 3D space. Based on this skeletal capture data, the content provision system 1 can determine the actions of the performer 19, such as dancing, posing, or specific postures. In addition to the performer 19, the system may also determine the movements of microphones, musical instruments, stage sets, etc., used by the performer 19.

[0018] In recent years, in the field of sports such as soccer and basketball, a technology called EPTS (Electronic Performance and Tracking Systems) has become known for estimating the posture and position of players and referees, as well as the position and rotation of the ball, from a designated field, using images from specially installed cameras and information from sensors (accelerometers and GPS sensors) attached to people (players) and objects (balls) involved in the game.

[0019] Specifically, camera 6 may be configured to capture images to obtain such EPTS data as skeletal capture data. This allows for precise determination of the performer's actions 19. The images captured by camera 6 can also be used as live footage of events such as music concerts.

[0020] At least one PTZ camera 7 is installed and is capable of imaging the performers 19 on stage 16. The PTZ camera 7 is a camera that can change its imaging direction in the pan and tilt directions, as well as change its zoom magnification (angle of view), based on control signals CT from the server device 2. This PTZ camera is mainly controlled to appropriately capture images of the performers 19's faces and bodies.

[0021] The PTZ camera 7 may be mounted on a tripod or the like, allowing for variable pan and tilt directions, or the lens barrel may be displaced relative to the camera body, allowing for variable pan and tilt directions.

[0022] The local controller 4 is shown as a device that performs communication processing at the live venue 15. For example, the captured image MCVs captured by each camera 6 and the PTZ camera 7 are supplied to the local controller 4 and transmitted to the server device 2 by the local controller 4. Alternatively, each camera 6 may be equipped with a communication function or connected to a corresponding communication device to directly transmit the captured image MCVs to the server device 2. In that case, the local controller 4 may not be necessary.

[0023] The player device 5 is an information processing device such as a smartphone, tablet terminal, or personal computer, but this player device 5 is assumed to be a device owned by a general user. The display unit of the player device 5 displays content images and the like provided by the demonstration collaboration service.

[0024] Server device 2 performs various processes for distributing content to player device 5 as part of the demonstration collaboration service. For example, as a processor indicated as the calculation unit 10, server device 2 performs processes such as analyzing the MCV of images and audio captured by camera 6 and PTZ camera 7 to generate motion data of the performer 19, predicting posture, controlling PTZ camera 7, personal identification, and integration processing that links motion data with person identification information.

[0025] This server device 2 can be implemented as an information processing device that performs cloud computing, i.e., a cloud server, or it can be implemented by a computer device installed at the live venue 15.

[0026] DB section 3 comprehensively shows the storage that the server device 2 uses to store various databases for the demonstration collaboration service. In this embodiment, a motion database for posture prediction processing and a person database for person identification may be provided.

[0027] An example of how image content can be provided by such a content provision system 1 is shown. For example, suppose a group of singers is performing on stage 16. Multiple cameras 6 capture the performance from various angles, and the MCV images from each camera 6 are transmitted to the server device 2. The server device 2 uses the MCV images to perform skeletal capture of each performer 19 in three-dimensional space and can determine their posture and movement, which is a change in posture.

[0028] Figure 3 schematically shows motion data MT1. In three-dimensional space, the body's central position coordinates and the position coordinates of each joint are determined for each performer 19 and converted into motion data MT1. This makes it possible to recognize the position and posture of each performer 19 for each frame of the captured image.

[0029] On the other hand, the content provision system 1 associates each performer 19 with each CG character on a one-to-one basis. That is, a CG character is prepared for each performer 19. Then, using the motion data MT1 of each performer 19, the system determines the position and posture of each performer 19 and generates an image in which that performer 19 is replaced by the CG character. Figure 4 shows an example of a content image using the CG character 30. In other words, this is an example of an image in which the performance of a real performer 19 is replaced by the CG character 30. To put it another way, it can be said that this is content that reproduces the performance of the performer 19 using the CG character 30.

[0030] Therefore, an image is generated in which the position and posture of the CG character 30 are set based on the motion data MT1 of each performer 19 in each frame of the captured image MCV. However, for this to work, it is necessary to be able to determine which performer 19's position and posture each motion data MT1 represents, as shown in Figure 3. Therefore, person recognition, such as facial recognition, is performed on each performer 19 on the stage to identify individuals, and the person identification information IDp is associated with the motion data MT1.

[0031] In the system, as shown in Figure 5, each of the 19 performers is assigned a unique person identification information IDp, such as "P001" or "P002," and these person identification information IDp are associated with each CG character 30, which is uniquely associated with character identification information IDc, such as "C001" or "C002." In other words, a relationship is established where "this performer is this character."

[0032] As a result, the person identification information IDp and motion data MT1 are associated, and the CG character 30 (character identification information IDc) whose position and posture are set by the motion data MT1 is identified. Therefore, the server device 2 can generate and provide content images like those in Figure 4 by generating integrated data MT2 for each performer 19, which is an integrated combination of motion data MT1 and person identification information IDp.

[0033] The server device 2 only needs to generate the integrated data MT2, and the process of generating content images using the integrated data MT2 may be performed by a subsequent information processing device, such as the player device 5. Alternatively, an information processing device that acts as a content generation device may be interposed between the server device 2 and the player device 5 in Figure 1.

[0034] <2. Server Equipment Configuration>

[0035] Figure 6 shows an example configuration of an information processing device 70 that functions as a server device 2. The information processing device 70 can be configured as, for example, a dedicated workstation, a general-purpose personal computer, a mobile terminal device, etc.

[0036] The CPU 71 of the information processing device 70 shown in Figure 6 executes various processes according to programs stored in the ROM 72 or non-volatile memory section 74 such as EEP-ROM (Electrically Erasable Programmable Read-Only Memory), or programs loaded from the storage section 79 into the RAM 73. The RAM 73 also appropriately stores data necessary for the CPU 71 to execute various processes.

[0037] The image processing unit 85 is configured as a processor that performs various image processing tasks. For example, it is a processor that can perform any of the following: generating animated images or CG images, image analysis processing on captured images, image effect processing, image generation processing such as video clips, or DB (Database) processing.

[0038] This image processing unit 85 can be implemented, for example, by a separate CPU, GPU (Graphics Processing Unit), GPGPU (General-purpose computing on graphics processing units), AI (artificial intelligence) processor, etc., distinct from the CPU 71. Alternatively, the image processing unit 85 may be provided as a function within the CPU 71.

[0039] The CPU 71, ROM 72, RAM 73, non-volatile memory unit 74, and image processing unit 85 are interconnected via a bus 83. An input / output interface 75 is also connected to this bus 83.

[0040] An input unit 76, consisting of operators or operating devices, is connected to the input / output interface 75. For example, the input unit 76 can be various operators or operating devices such as a keyboard, mouse, keys, dial, touch panel, touchpad, or remote controller. User operations are detected by the input unit 76, and the signals corresponding to the input operations are interpreted by the CPU 71.

[0041] Furthermore, a display unit 77, consisting of an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, and an audio output unit 78, consisting of a speaker, are connected to the input / output interface 75, either as an integrated unit or as separate components.

[0042] The display unit 77 performs various displays as a user interface. The display unit 77 is composed of, for example, a display device provided on the housing of the information processing apparatus 70, a separate display device connected to the information processing apparatus 70, or the like. The display unit 77 executes various image displays on the display screen based on the instructions of the CPU 71. Further, the display unit 77 performs displays such as various operation menus, icons, messages, etc., that is, displays as a GUI (Graphical User Interface) based on the instructions of the CPU 71.

[0043] The input / output interface 75 may be connected to a storage unit 79 composed of an SSD (Solid State Drive), an HDD (Hard Disk Drive), etc., and a communication unit 80 composed of a modem, etc. The communication unit 80 performs communication processing via a transmission path such as the Internet, wired / wireless communication with various devices, bus communication, etc.

[0044] The input / output interface 75 is also connected to a drive 82 as necessary, and a removable recording medium 81 such as a flash memory, a memory card, a magnetic disk, an optical disk, a magneto-optical disk, etc. is appropriately mounted. The drive 82 can read data files such as image files and various computer programs from the removable recording medium 81. The read data files are stored in the storage unit 79, or images and sounds included in the data files are output by the display unit 77 and the audio output unit 78. Further, computer programs etc. read from the removable recording medium 81 are installed in the storage unit 79 as necessary.

[0045] In this information processing apparatus 70, software can be installed via network communication by the communication unit 80 or via the removable recording medium 81. Alternatively, the software may be stored in the ROM 72, the storage unit 79, etc. in advance.

[0046] For example, when considering this information processing apparatus 70 as the server apparatus 2, a function for performing the processes of FIGS. 7 and 9 described later by software is provided, and the processes as the arithmetic unit 10 by those functions are executed by the CPU 71 and the image processing unit 85.

[0047] <3. Processing Example> The following describes a processing example by the arithmetic unit 10 of the server apparatus 2. FIG. 7 shows a processing example executed by the arithmetic unit 10 during a live performance by the performer 19, for example.

[0048] During the live performance, the captured image MCV of the live venue 15 by the camera 6 and the PTZ camera 7 is continuously transmitted to the server apparatus 2. The arithmetic unit 10 performs analysis processing on the captured image and the audio accompanying the captured image in real time, and executes the process of FIG. 7.

[0049] FIG. 7 shows a motion data generation process S1, a posture determination process S2, a posture prediction process S3, a control signal generation process S4, a person identification process S5, and an integration process S6. These can be regarded as the processing functions of the arithmetic unit 10 realized by a software program.

[0050] The arithmetic unit 10 performs the motion data generation process S1 for each frame of the input captured image MCV, particularly the captured images MCV6 of the plurality of cameras 6. This is a process of generating the motion data MT1 of each performer 19 in the three-dimensional space recognizable by the captured images MCV6 by the plurality of cameras 6. As described above, the motion data MT1 is information indicating the body center coordinates and the coordinates of each joint position of the performer 19 in the actual three-dimensional space including the stage 16, for example. The arithmetic unit 10 generates such motion data MT1 for each frame timing (or for each intermittent frame timing) of the captured image MCV6 and for each performer 19. By using the captured images MCV6 of the plurality of cameras 6, the joint positions and the like can be represented as XYZ coordinate values in the three-dimensional space, and such motion data MT1 becomes skeleton capture data corresponding to the position and posture of each performer 19 in the three-dimensional space.

[0051] This motion data MT1 is output as integrated data MT2 via integration processing S6. In integration processing S6, if person identification information IDp is input by person identification processing S5 for a certain performer 19, processing is performed to associate that person identification information IDp with the corresponding motion data MT1.

[0052] In motion data generation process S1, tracking is performed by assigning tracking identification information IDt to each performer 19. As the frames of the captured image MCV progress, motion data MT1 with the same tracking identification information IDt are recognized as belonging to the same person.

[0053] In other words, when a performer 19 is recognized in the image, tracking identification information IDt is set, and motion data MT1 is generated. Since the motion data MT1 of each performer in each frame is associated with the tracking identification information IDt, as long as tracking is maintained, it is possible to recognize which performer 19's position and posture each motion data MT1 generated in each frame represents in relation to the preceding and succeeding frames.

[0054] Figure 8 shows the information managed by tracking identification information IDt. For each tracking identification information IDt (for example, "T001"), motion data MT1 of the performer 19 is generated, and at some point, the performer 19's person identification information IDp is associated with it. During the generation of motion data MT1, the motion score is updated as an evaluation value for the reliability (3D position accuracy) of the motion data MT1. The tracking score is also updated as an evaluation value for the accuracy of the tracking itself. Therefore, in the processing process of the calculation unit 10, as shown in the figure, the motion data MT1, person identification information IDp, motion score, and tracking score are associated with the tracking identification information IDt.

[0055] The person identification information IDp is associated with the tracking identification information IDt in the integration process S6, which occurs at the time the person identification information IDp is obtained in the person identification process S5. This associates it with the motion data MT1 as shown in Figure 8. The integrated data MT2 is output with at least the motion data MT1 and the person identification information IDp associated with it.

[0056] For each performer 19, at least motion data MT1 and person identification information IDp are output as integrated data MT2, so that in subsequent processing, i.e., processing of information processing devices other than the calculation unit 10 or server device 2, an animated image of CG characters 30 as shown in Figure 4 can be generated. As shown in Figure 5, since the person identification information IDp and character identification information IDc are linked in the system, the CG character 30 to which motion data MT1 is applied can be determined by the person identification information IDp. In other words, the position and posture of each CG character 30 can be determined frame by frame by the motion data MT1 of each performer 19. Therefore, by outputting integrated data MT2 frame by frame, an animated image of CG characters 30 that reproduces the performance of each performer 19 can be generated.

[0057] As can be understood from the above processing, in order to provide image content using CG character 30 that reproduces the performance, it is necessary that the person identification information IDp is correctly associated with the motion data MT1. However, in cases such as singing accompanied by dance performances where the movements are vigorous and the positions of the members change significantly, person identification and tracking within the captured image MCV is difficult.

[0058] Therefore, the calculation unit 10 sequentially performs posture determination processing S2, posture prediction processing S3, control signal generation processing S4, and person identification processing S5, and based on this, integration processing S6 is performed. This involves selecting motion data MT1 (i.e., a certain performer 19) with a certain tracking identification information IDt, performing person identification on that performer 19, and setting or updating person identification information IDp corresponding to the motion data MT1.

[0059] In the posture determination process S2, the calculation unit 10 determines the posture of a performer 19 based on the motion data MT1 obtained in the motion data generation process S1. In other words, it determines the posture of the performer 19 indicated by a certain tracking identification information IDt. Then, in the posture prediction process S3, the calculation unit 10 predicts the posture and position of the performer 19 after a certain time has elapsed, that is, the posture and position after the present moment, based on the performer 19's current posture. Here, the later time is the time after which the PTZ camera 7 has sufficient time to capture the person's face and body. This time can be a fixed amount of time, but it does not necessarily have to be a fixed amount of time.

[0060] For example, the calculation unit 10 predicts the position and posture of the performer 19 on the stage 16 after a set period of time, for example, 2 to 5 seconds, taking into account the pan, tilt, and zoom operations of the PTZ camera 7. It is particularly useful to predict the position and orientation of the face. Alternatively, the calculation unit 10 may predict whether the position and orientation of the performer 19's body and face will be suitable for imaging by the PTZ camera 7 within a time period, for example, within 10 seconds from the current frame.

[0061] In the posture prediction process S3, the calculation unit 10 can predict the posture using linear prediction. For example, there is a method to predict the position and posture a few seconds later using linear prediction from the position and posture of the current frame and past frames.

[0062] The calculation unit 10 may also refer to the motion DB 3a in the posture prediction process S3. For example, the motion DB 3a may be a database of learning data on human movement, specifically a database of learning data on various human movements, particularly dance performance movements. Specifically, it is information for determining what position and posture a person is likely to be in a few seconds, based on posture determination. The calculation unit 10 can perform learning-based prediction processing by referring to the motion DB 3a.

[0063] Furthermore, if the content of the performance, such as choreography to match the music, is known in advance, the motion DB 3a may store the progress data of that choreography. By referring to this data, the calculation unit 10 can predict the posture and position of the performer 19 from the present moment onward.

[0064] Based on the results of the posture prediction process S3, the position and posture of the performer 19 on stage a few seconds from the present, and in particular the position and orientation of their face based on that posture, are predicted. Based on this information, the calculation unit 10 performs the control signal generation process S4.

[0065] Specifically, control signals CT for pan, tilt, and zoom (part or all) to the PTZ camera 7 are generated and output, in accordance with the predicted orientation and position of the performer's face 19, so that the face can be properly imaged. The control signals CT are transmitted to the PTZ camera 7, for example, via the local controller 4, and some or all of the pan, tilt, and zoom are executed in the PTZ camera 7. This greatly increases the possibility of capturing a close-up image of the performer's face 19 immediately afterward. Alternatively, the PTZ camera 7 may capture the entire body of the performer 19 instead of just a close-up of the face. In any case, the goal is to capture an image with the performer 19 as the main subject in the PTZ camera 7.

[0066] After outputting the control signal CT in the control signal generation process S4, the calculation unit 10 performs the person identification process S5. In this case, the calculation unit 10 analyzes the facial image of a person using the captured image MCV, particularly the captured image MCV7 from the PTZ camera 7. In this case, the calculation unit 10 refers to the person DB 3b. The person DB 3b stores information about the face and height of each performer 19. The calculation unit 10 identifies individuals by comparing the facial features of the person analyzed from the captured image MCV7 with the facial feature information stored in the person DB 3b. In addition to facial features, other features such as height and skeletal structure may also be compared.

[0067] When the calculation unit 10 identifies an individual, it provides the person identification information IDp assigned to that person to the integration process S6. As a result, the person identification information IDp has been determined for performer 19 to whom a certain tracking identification information IDt has been assigned. Therefore, in the integration process S6, the calculation unit 10 associates the person identification information IDp with the corresponding tracking identification information IDt. As a result, as shown in Figure 8, the motion data MT1 and the person identification information IDp have been associated with performer 19 to whom a certain tracking identification information IDt has been assigned.

[0068] In this way, once the person identification information IDp is associated with the tracking identification information IDt, the relationship between the tracking identification information IDt and the person identification information IDp is maintained in subsequent frames. Therefore, unless the tracking identification information IDt is lost due to reasons such as the loss of tracking for the performer 19, the relationship between the tracking identification information IDt and the person identification information IDp will be maintained. Motion data MT1 is generated for each frame, but because the relationship between the tracking identification information IDt and the person identification information IDp is maintained, the association of the person identification information IDp with the motion data MT1 in each frame is maintained.

[0069] The posture determination process S2, posture prediction process S3, control signal generation process S4, person identification process S5, and integration process S6 (hereinafter collectively referred to as the "identification integration process") described above are performed sequentially for each performer 19 (each tracking identification information IDt). As a result, the server device 2 can output integrated data MT2 for all performers 19, which links motion data MT1 with person identification information IDp. Based on the integrated data MT2 as described above, the server device 2 or the player device 5 can generate animated images that reproduce the performance by the CG character 30.

[0070] Next, we will explain the priority determination for executing the identification and integration processes (S2, S3, S4, S5, S6) described above. By performing the identification and integration process once for a certain performer 19 (a certain tracking identification information IDt), the relationship between the tracking identification information IDt and the person identification information IDp is maintained thereafter. However, since it is not always possible to maintain the relationship with absolute accuracy, it is appropriate to perform the identification and integration process for each performer 19 sequentially and update the relationship between the tracking identification information IDt and the person identification information IDp. Therefore, it is desirable to perform the identification and integration process for each performer 19 at a somewhat regular pace.

[0071] Furthermore, tracking of performer 19 is not always accurate. Tracking accuracy may decrease if performers cross paths and change positions during a performance, or if they adopt unusual postures. In some cases, the tracking identification information IDt for one performer 19 may cause another performer 19 to be tracked midway through the performance. In addition, if a performer 19 is hidden by the stage set or goes out of frame, tracking may become impossible, and the tracking identification information IDt may be lost during the tracking process. Moreover, the accuracy of posture prediction or unexpected movements of the performer 19 may reduce the accuracy of person identification from the image MCV7 captured by the PTZ camera 7.

[0072] Given these circumstances, for example, the calculation unit 10 determines the performer 19, specifically the tracking identification information IDt, to be targeted for identification and integration processing in the process shown in Figure 9. The calculation unit 10 repeatedly executes the process shown in Figure 9 during the performance, and each time it determines the target performer 19 (tracking identification information IDt), it executes the identification and integration processing described in Figure 7.

[0073] In step S101, the calculation unit 10 checks whether new tracking identification information IDt has been generated. That is, it checks whether a performer 19 that was not tracked up to the previous frame has been detected in the captured image MCV6, and whether a new tracking identification information IDt has been set and tracking has started. This includes not only when a new performer 19 appears on stage 16, but also when tracking is restarted for a performer 19 that had previously lost tracking.

[0074] In these cases, since a person identification information IDp has not yet been associated with the new tracking identification information IDt, the identification integration process is given priority. The calculation unit 10 proceeds to step S120, selects one of the new tracking identification information IDts, and in step S131, sets the selected tracking identification information IDt to be the target for the identification integration process. As a result, the identification integration process described in Figure 7 is performed on the newly generated tracking identification information IDt, and the person identification information IDp is associated with the tracking identification information IDt by the integration process S6.

[0075] If multiple performers 19 appear on stage 16 at once, steps S101 and S120 will be performed sequentially for each performer 19. Therefore, the existence of tracking identification information IDt that is not associated with person identification information IDp for a long period of time is avoided.

[0076] If no new tracking identification information IDt is generated, the calculation unit 10 proceeds to step S102 to check whether there are any tracking identification information IDt for which identification integration processing has not been performed for a threshold time thT or longer. It is desirable to perform identification integration processing for each performer 19 at certain time intervals. Therefore, a threshold time thT is set, and if there are performers 19 for whom identification integration processing has not been performed beyond the threshold time thT, those performers 19 are given priority.

[0077] If there are performers 19 who have not undergone the identification and integration process beyond the threshold time thT, the calculation unit 10 proceeds to step S121, selects one of the relevant tracking identification information IDts, and in step S131 sets the selected tracking identification information IDt to be the target for the identification and integration process.

[0078] As a result, the identification and integration process described in Figure 7 is performed for performers 19 (tracking identification information IDt) whose identification and integration processes have been performed at intervals, and the person identification information IDp is updated for that tracking identification information IDt by the integration process S6.

[0079] If no performers 19 have not undergone identification and integration processing during the threshold time thT, the calculation unit 10 proceeds to step S103 to check whether there is any tracking identification information IDt whose tracking score is less than or equal to the threshold thS1. For tracking processing, the calculation unit 10 scores the tracking accuracy for each performer 19. For example, if there is an event that reduces tracking accuracy, such as the intersection or overlap of the positions of the performers 19, the tracking score is lowered. For tracking identification information IDt whose tracking score has decreased, it is better to redo the identification and integration processing.

[0080] If a corresponding tracking identification information IDt exists, the calculation unit 10 proceeds to step S122 and updates the identification execution score for tracking identification information IDt whose tracking score is less than or equal to the threshold thS1. The identification execution score is a score used to determine the priority of the identification integration process. For example, if the tracking scores of tracking identification information IDt = T001 and T005 are less than or equal to the threshold thS1, the identification execution score of tracking identification information IDt = T001 and T005 is incremented to increase their priority.

[0081] After processing the tracking score, the calculation unit 10 proceeds from step S103 or step S122 to step S104 to check whether there is tracking identification information IDt whose motion score is less than or equal to the threshold thM. In the motion data generation process S1, the calculation unit 10 generates motion data MT1 for each performer 19, but the reliability of the motion data MT1 may not be high depending on the posture and position of the performer 19. For example, joint position detection is difficult in lying down, squatting, or extreme postures. In such cases, the detection accuracy of the motion data MT1 may decrease. Therefore, the calculation unit 10 sets a motion score as information indicating the reliability of the motion data MT1. That is, the motion score is lowered when there is an event that reduces the reliability of the motion data MT1.

[0082] Therefore, in step S104, the motion score is determined, and if there is a tracking identification information IDt whose motion score is less than or equal to the threshold thM, the calculation unit 10 proceeds to step S123 and updates the identification execution score for the corresponding tracking identification information IDt.

[0083] There are two approaches to motion scoring. For performers 19 with a low motion score, tracking accuracy decreases. If this is the focus, the identification execution score value for tracking identification information IDt with a low motion score is increased, and its priority is raised. On the other hand, for performers 19 with a low motion score, the accuracy of posture prediction processing S3 decreases, and as a result, the accuracy of person identification processing S5 decreases. If this is the focus, the identification execution score value for tracking identification information IDt with a low motion score is lowered, and its priority is reduced.

[0084] After processing the motion score, the calculation unit 10 proceeds from step S104 or step S123 to step S105 to check whether there is tracking identification information IDt with a positional advantage. Positional advantage means that the performer 19 is in a position suitable for imaging by the PTZ camera 7.

[0085] For example, if multiple performers 19 are lined up in multiple rows from the front to the back on the stage 16, the performers 19 in the front row are considered to have a positional advantage. The performers 19 in the back row are considered to have no positional advantage, as they may be hidden by the performers 19 in the front row from the perspective of the PTZ camera 7. There is an advantage or disadvantage in terms of imaging by the PTZ camera 7 for each position on the stage 16. And, regarding the identification and integration processing, it is expected that the accuracy of person identification will be improved by performing it on the performers 19 who have a positional advantage.

[0086] The calculation unit 10 then determines the existence of a performer 19 with a positional advantage based on the position of each performer 19 in three-dimensional space and their relative positions to one another. If a performer 19 with a positional advantage (tracking identification information IDt) exists, the calculation unit 10 proceeds to step S124 and updates the identification execution score for the corresponding tracking identification information IDt. In this case, the identification execution score is updated to increase the priority.

[0087] After processing for positional advantage, the calculation unit 10 proceeds from step S105 or step S124 to step S130 and selects one tracking identification information IDt based on the identification execution score for each tracking identification information IDt. Then, in step S131, the selected tracking identification information IDt is set as the target of the identification integration process.

[0088] As a result, the tracking identification information IDt selected from the perspectives of tracking score, motion score, and positional advantage undergoes the identification integration process described in Figure 7, and the person identification information IDp is associated with the tracking identification information IDt by the integration process S6.

[0089] As shown in Figure 9 above, the target for the identification and integration process is selected, and the association between motion data MT1 and person identification information IDp is appropriately updated for each performer 19, improving the accuracy of the association.

[0090] <4. Summary and Modifications> The above embodiments provide the following effects.

[0091] The information processing device, which serves as the server device 2 in this embodiment, includes a calculation unit 10 that performs posture prediction processing S3, which analyzes the captured image MCV of the performer 19 as the subject to predict the position and posture of the performer 19 at a later point in time, and control signal generation processing S4, which generates a control signal CT that changes the imaging direction or zoom magnification of the PTZ camera 7 with the performer 19 as the target, according to the prediction result of posture prediction processing S3. In other words, it estimates the posture of the performer 19 at a later point in time, and changes the imaging direction and zoom magnification of the PTZ camera 7 so that the performer 19's face can be appropriately captured at the time when that posture is achieved. This makes it possible to perform PTZ control that matches the timing when the orientation of the performer 19's face is appropriate for the PTZ camera 7. Therefore, the performer 19's face can be appropriately captured by the PTZ camera 7.

[0092] In the posture prediction process S3, the position and posture of the performer 19 are predicted, but it is also possible to predict only the position. For example, in cases where person identification is performed from the entire body, the position prediction allows the PTZ camera 7 to be controlled for the target performer 19, and the performer 19's body can be imaged. Alternatively, the posture prediction process S3 may predict only the posture. For example, in performances where there is little change in position, posture prediction may be performed to accurately determine the position of the face, etc.

[0093] Although an example using one PTZ camera 7 was given, multiple PTZ cameras 7 may be used, and the identification and integration processing may be performed in parallel for multiple performers 19. In that case, the control signal generation processing S4 should be performed for each PTZ camera 7 according to the attitude prediction processing S3. Although a PTZ camera 7 was given as an example, a camera with variable pan direction only, a camera with variable tilt direction only, a camera with a fixed imaging direction and variable zoom magnification only, or a camera with variable pan and tilt directions and a fixed zoom magnification may be used.

[0094] In this embodiment, the calculation unit 10 predicts the posture of the performer 19 from the present moment onward in posture prediction processing S3 using posture determination results obtained by analyzing the captured images MCV6 from multiple cameras 6 that image the performer 19, and generates a control signal CT for the PTZ camera 7, which is capable of controlling changes in imaging direction or zoom magnification, in control signal generation processing S4. For example, the entire live stage is imaged from various angles using a multi-camera system with multiple first imaging devices (cameras 6), and the posture is determined three-dimensionally. From the determined posture, the posture after a predetermined time is predicted using linear prediction, human movement learning-based prediction, etc. This allows for setting appropriate control target direction and zoom magnification in control signal generation processing S4. Furthermore, for person identification, it is desirable to capture close-up images of the body or face of a specific performer 19. Therefore, a control signal CT is generated for the PTZ camera 7, which is a second imaging device. This makes it possible to obtain captured images suitable for person identification as the captured images MCV7 from the PTZ camera 7.

[0095] In this embodiment, the calculation unit 10 generates a control signal CT targeting the face of the performer 19 in the control signal generation process S4. Considering that person identification is performed using facial images, it is preferable to perform pan, tilt, and zoom control targeting the face of the performer 19 to capture a larger image of the face, in order to improve identification accuracy.

[0096] In this embodiment, the calculation unit 10 analyzes the captured images MCV6 from multiple cameras 6 and performs motion data generation processing S1 to generate motion data MT1 that indicates the movement or posture of each performer 19 appearing in the captured images MCV6. The entire live stage is captured from various angles using a multi-camera system with multiple first imaging devices (cameras 6), and a process is performed to estimate the posture in three dimensions. This is suitable for determining the posture of multiple performers 19 and generating motion data MT1.

[0097] In this embodiment, the calculation unit 10 performs a person identification process S5 that analyzes the captured image MCV7 from the PTZ camera 7 and generates person identification information IDp. By appropriately controlling the PTZ camera 7, which is the second imaging device, the probability of the captured image MCV7 from the PTZ camera 7 being an image of a performer's face or body in a close-up and easily viewable orientation increases. By using such an captured image MCV7, the reliability of the person identification process S5, which identifies individual people, can be increased.

[0098] In this embodiment, the calculation unit 10 performs an integration process S6 that associates the person identification information IDp generated in the person identification process S5 with the motion data MT1 of the corresponding performer 19. This allows the motion data MT1 associated with the person identification information IDp to be output as integrated data MT2, enabling subsequent processing to identify each performer 19 and perform image processing according to the posture and movement of each performer 19. For example, the player device 5 can generate images in which a CG character 30 associated with each performer 19 performs the dance performance of multiple performers 19.

[0099] In this embodiment, the calculation unit 10 analyzes the captured images MCV6 from multiple cameras 6, tracks each performer 19 appearing in the captured images MCV6, manages them with tracking identification information IDt, and then performs motion data generation processing S1 to generate motion data MT1 that indicates the movement or posture of the performer 19. Each of the multiple performers 19 is tracked in a three-dimensional space recognized, for example, by multi-camera images, and managed with tracking identification information IDt. Motion data MT1 is associated with the tracking identification information IDt in each frame, and further associated with person identification information IDp. As a result, the association between motion data MT1 and person identification information IDp is maintained as long as tracking is maintained. Therefore, the person identification processing S5 and the integration processing S6 that associates person identification information IDp with motion data MT1 do not need to be performed continuously for each performer 19, but can be performed intermittently. This also reduces the system processing burden.

[0100] In this embodiment, the calculation unit 10 updates the person identification information IDp associated with the motion data MT1 managed by the tracking identification information IDt of the performer 19 who was the target of the person identification process S5 during the integration process S6, in response to the person identification process S5 being performed. As a result, the association between the motion data MT1 managed by the tracking identification information IDt and the person identification information IDp is updated sequentially, allowing the latest identification state to be achieved. For example, even if tracking is lost, the association between the motion data MT1 and the person identification information IDp can be restored to the correct state.

[0101] In this embodiment, the calculation unit 10 performs a priority determination on the performer 19 appearing in the captured image MCV6 of the camera 6 to select the performer 19, and then performs posture prediction processing S3, control signal generation processing S4, and person identification processing S5. By appropriately selecting the performer 19 that generates person identification information IDp through priority determination, the identification integration processing (S2, S3, S4, S5, S6) can be performed efficiently, and the association between motion data MT1 and person identification information IDp can be maintained in a correct state.

[0102] The priority determination in this embodiment includes a process to increase the priority of the performer 19 whose tracking has just been started (see steps S101 and S120 in Figure 9). For a performer whose tracking has just been started, the motion data MT1 and the person identification information IDp are not yet associated. Therefore, it is preferable to execute the person identification process S5 with priority to quickly associate the motion data MT1 and the person identification information IDp.

[0103] The priority determination in this embodiment includes a process to increase the priority of performers 19 for whom the person identification process S5 has not been performed for a predetermined time or longer (see steps S102 and S121 in Figure 9). When the identification and integration process is performed intermittently for each performer 19, if the person identification process S5 has not been performed for a long time, the accuracy of the integrated data MT2 may decrease. Therefore, it is preferable to prioritize the execution of the identification and integration process for performers 19 for whom the person identification process S5 has not been performed for a certain amount of time or longer.

[0104] The priority determination in this embodiment includes a process to increase the priority of performers 19 whose tracking status evaluation value is below a predetermined level (see steps S103 and S122 in Figure 9). In the tracking process, the tracking score is updated sequentially as an evaluation value indicating the accuracy of tracking. A low tracking score means that the reliability of tracking for performers 19 has decreased, and therefore the reliability of the correspondence between motion data MT1 and person identification information IDp has decreased. For this reason, it is preferable to prioritize the identification and integration process for the performer 19 in question.

[0105] The priority determination in this embodiment includes a process to change the priority of performers 19 whose reliability evaluation value of motion data MT1 is below a predetermined level (see steps S104 and S123 in Figure 9). The motion score indicating the reliability of motion data MT1 is updated sequentially. If the motion score is low and the reliability of motion data MT1 is reduced, the priority is either increased or decreased. This allows the reliability of motion data MT1 to be reflected in the priority of the identification and integration process.

[0106] The priority determination in this embodiment includes a process to increase the priority of performers 19 who are determined to have a positional advantage based on their location information (see steps S105 and S124 in Figure 9). For example, in a dance performance, performers 19 who are in the front row, or performers 19 who are in a position where their faces can be easily captured by the PTZ camera 7, are determined to have a positional advantage in terms of person identification. It is expected that these performers 19 can be identified with high accuracy from the images captured by the PTZ camera 7. Therefore, it is preferable to prioritize the execution of the person identification process S5 for these performers 19.

[0107] In the content provision system 1 of the embodiment, an example was described in which image content is provided that reproduces the performance of the performer 19 using a CG character 30. However, the technology of this disclosure can be applied to a variety of processes. For example, by obtaining integrated data MT2 in which motion data MT1 and person identification information IDp for the performer 19 are associated, various diverse image processing becomes possible.

[0108] For example, in a live image of performer 19, it is possible to generate image content that incorporates special effects depending on the specific posture and movement of a particular performer 19. It is also possible to generate live images with different image effects for each performer 19. Furthermore, it is possible to generate virtual live images with different image effects for each character corresponding to performer 19.

[0109] Furthermore, it is possible to implement a service that distributes still images and video content as a bonus to the audience or viewers based on the movements of specific performers 19. It can also be applied to a system that automatically controls live performance elements such as lighting, smoke, special effects, set changes, and background images in accordance with the movements and postures of specific performers 19.

[0110] Furthermore, the attitude prediction processing S3 and control signal generation processing S4 of this technology can also be applied as camera control to appropriately capture a specific performer 19 when incorporating images captured by the PTZ camera 7 into content images.

[0111] In this embodiment, the server device 2 is configured as a cloud server, but it may also be configured as a local server installed at a live venue 15 or the like.

[0112] The technology of the content provision system 1 disclosed herein can be applied not only to live music performances but also to various other performances such as entertainment and sports. For example, performers 19 can include singers, actors, entertainers, dancers, musicians, athletes, orators, speakers, magicians, and others who perform some kind of entertainment act in front of an audience. This technology can be applied when it is desired to associate the motion data MT1 of the performer 19 with person identification information IDp during the performance opportunities of these performers 19.

[0113] Furthermore, the performers are not limited to real people. They can be animals, virtual humans, or characters. For example, this can be applied to performances in the metaverse space.

[0114] The program of the embodiment is a program that causes an information processing device 70, such as a CPU, DSP (digital signal processor), AI processor, etc., or including these, to execute the processes described in Figures 7 and 9. Specifically, the program of the embodiment is a program that causes an information processing device to execute a posture prediction process S3 that analyzes an captured image MCV of the performer 19 as the subject and predicts the position or posture of the performer 19 from the present moment onward, and a control signal generation process S4 that generates a control signal CT that changes the imaging direction or zoom magnification of the imaging device with the performer 19 as the target, according to the prediction result of the posture prediction process S3.

[0115] With such a program, the information processing device 70, which will be the server device 2 of the embodiment, can be realized in, for example, a computer device, a mobile terminal device, or other device capable of performing information processing.

[0116] Such programs can be pre-recorded on HDDs (hard disk drives) or ROMs (remote-controlled memory) within microcomputers, which are recording media built into computer devices. Alternatively, programs can be temporarily or permanently stored (recorded) on removable recording media such as flexible disks, CD-ROMs (Compact Disc Read Only Memory), MO (Magneto Optical) disks, DVDs (Digital Versatile Discs), Blu-ray Discs (registered trademark), magnetic disks, semiconductor memory, and memory cards. Such removable recording media can be provided as so-called packaged software. Furthermore, such programs can be installed from removable recording media to personal computers, etc., or downloaded from download sites via networks such as LANs (Local Area Networks) and the Internet.

[0117] Furthermore, such a program is suitable for providing the information processing device 70 that constitutes the server device 2 of the embodiment to a wide range of users. For example, by downloading the program to mobile terminal devices such as smartphones and tablets, imaging devices, mobile phones, personal computers, game consoles, video equipment, PDAs (Personal Digital Assistants), etc., these devices can be made to function as the information processing device 70 that constitutes the server device 2 of this disclosure.

[0118] Furthermore, the effects described herein are merely illustrative and not limited to those described herein, and other effects may also occur.

[0119] The technology can also be configured as follows: (1) An information processing device comprising a calculation unit that performs: posture prediction processing, which analyzes captured images of a performer as the subject to predict the performer's position or posture after the present moment; and control signal generation processing, which generates a control signal to change the imaging direction or zoom magnification of the imaging device with the performer as the target, according to the prediction result of the posture prediction processing. (2) The information processing device according to (1) above, wherein the calculation unit predicts the performer's position or posture after the present moment using posture determination results obtained by analyzing captured images of a plurality of first imaging devices that image the performer, in the posture prediction processing, and generates a control signal for a second imaging device that is capable of changing the imaging direction or zoom magnification, in the control signal generation processing. (3) The information processing device according to (1) or (2) above, wherein the calculation unit generates the control signal with the performer's face as the target, in the control signal generation processing. (4) The information processing device according to (2) above, wherein the calculation unit performs motion data generation processing, which analyzes captured images of a plurality of the first imaging devices to generate motion data indicating the movement or posture of each performer appearing in the captured images. (5) The information processing device according to (2) or (4) above, wherein the calculation unit performs a person identification process to analyze the captured image of the second imaging device and generate person identification information. (6) The information processing device according to (5) above, wherein the calculation unit performs a motion data generation process to analyze the captured images of a plurality of first imaging devices and generate motion data indicating the movement or posture of each performer appearing in the captured image, and integrates the identification information generated in the person identification process to the motion data of the corresponding performer. (7) The information processing device according to (6) above, wherein the calculation unit analyzes the captured images of a plurality of first imaging devices, tracks each performer appearing in the captured image and manages them with tracking identification information, and then performs the motion data generation process. (8) The information processing device according to (7) above, wherein the calculation unit updates the identification information associated with the motion data managed by the tracking identification information of the performer who was the target of the person identification process in the integration process, in response to the person identification process being performed.(9) The information processing apparatus according to (7) or (8) above, wherein the calculation unit performs priority determination on performers appearing in the captured image of the first imaging device to select a performer, and performs the posture prediction processing, the control signal generation processing, and the person identification processing. (10) The information processing apparatus according to (9) above, wherein the priority determination includes processing to increase the priority of performers whose tracking has just been started. (11) The information processing apparatus according to (9) or (10) above, wherein the priority determination includes processing to increase the priority of performers whose person identification processing has not been performed for a predetermined time or longer. (12) The information processing apparatus according to any one of (9) to (11) above, wherein the priority determination includes processing to increase the priority of performers whose tracking state evaluation value is below a predetermined level. (13) The information processing apparatus according to any one of (9) to (12) above, wherein the priority determination includes processing to change the priority of performers whose motion data reliability evaluation value is below a predetermined level. (14) The priority determination is an information processing device according to any one of (9) to (13) above, which includes a process to increase the priority of performers who are determined to have a positional advantage based on positional information. (15) An information processing method in which an information processing device performs a posture prediction process to predict the position or posture of a performer after the present time by analyzing an image captured with a performer as the subject, and a control signal generation process to generate a control signal that changes the imaging direction or zoom magnification of an imaging device with the performer's face as the target, according to the prediction result of the posture prediction process. (16) A program that causes an information processing device to perform a posture prediction process to predict the position or posture of a performer after the present time by analyzing an image captured with a performer as the subject, and a control signal generation process to generate a control signal that changes the imaging direction or zoom magnification of an imaging device with the performer's face as the target, according to the prediction result of the posture prediction process.

[0120] 1. Content provision system 2. Server device 3. DB unit 4. Local controller 5. Player device 6. Camera 7. PTZ camera 10. Processing unit 15. Live venue 16. Stage 19. Performer 30. CG character MT1. Motion data MT2. Integrated data IDp. Person identification information

Claims

1. An information processing device comprising a calculation unit that performs:

1. An attitude prediction process that analyzes captured images of a performer as the subject to predict the performer's position or attitude at a later point in time; and 2. A control signal generation process that generates a control signal to change the imaging direction or zoom magnification of the imaging device, targeting the performer, according to the prediction result of the attitude prediction process.

2. The information processing apparatus according to claim 1, wherein the calculation unit predicts the position or posture of the performer at a later time using posture determination results obtained by analyzing captured images from a plurality of first imaging devices that capture images of the performer in the posture prediction process, and generates a control signal for a second imaging device that is capable of changing the imaging direction or zoom magnification in the control signal generation process.

3. The information processing apparatus according to claim 1, wherein the calculation unit generates the control signal with the performer's face as the target in the control signal generation process.

4. The information processing apparatus according to claim 2, wherein the calculation unit analyzes images captured by a plurality of first imaging devices and performs motion data generation processing to generate motion data indicating the movement or posture of each performer appearing in the captured images.

5. The information processing apparatus according to claim 2, wherein the calculation unit performs a person identification process to analyze the captured image of the second imaging device and generate person identification information.

6. The information processing apparatus according to claim 5, wherein the calculation unit performs motion data generation processing, which analyzes images captured by a plurality of first imaging devices and generates motion data indicating the movement or posture of each performer appearing in the captured images, and integrates the identification information generated in the person identification processing with the corresponding performer's motion data.

7. The information processing apparatus according to claim 6, wherein the calculation unit analyzes images captured by a plurality of first imaging devices, tracks each performer appearing in the captured images, manages them with tracking identification information, and then performs the motion data generation process.

8. The information processing apparatus according to claim 7, wherein the calculation unit updates the identification information associated with the motion data managed by the tracking identification information of the performer who was the target of the person identification process in the integration process, in response to the person identification process being performed.

9. The information processing apparatus according to claim 7, wherein the calculation unit performs priority determination on performers appearing in the captured image of the first imaging device to select a performer, and performs the posture prediction processing, the control signal generation processing, and the person identification processing.

10. The information processing apparatus according to claim 9, wherein the priority determination includes a process to increase the priority of a performer whose tracking has just been initiated.

11. The information processing apparatus according to claim 9, wherein the priority determination includes a process to increase the priority of performers whose person identification process has not been performed for a predetermined time or longer.

12. The information processing apparatus according to claim 9, wherein the priority determination includes a process to increase the priority of performers whose tracking status evaluation value is below a predetermined level.

13. The information processing apparatus according to claim 9, wherein the priority determination includes a process for changing the priority of performers whose evaluation value of the reliability of motion data is below a predetermined level.

14. The information processing apparatus according to claim 9, wherein the priority determination includes a process to increase the priority of performers who are determined to have a locational advantage based on location information.

15. An information processing method comprising: an information processing device performing posture prediction processing, which analyzes an image captured with a performer as the subject to predict the performer's position or posture at a later point in time; and a control signal generation processing, which generates a control signal to change the imaging direction or zoom magnification of the imaging device with the performer's face as the target, according to the prediction result of the posture prediction processing.

16. A program that causes an information processing device to execute: an attitude prediction process that analyzes captured images of a performer as the subject to predict the performer's position or attitude at a later point in time; and a control signal generation process that generates a control signal to change the imaging direction or zoom magnification of the imaging device, targeting the performer's face, according to the prediction result of the attitude prediction process.

Citation Information

Patent Citations

  • Device, method, and program for retrieving person

    JP2010257450A

  • Camera system and camera control apparatus

    JP2014241578A

  • Information processing device, information processing program, information processing method, and information processing system

    JP2020173711A

  • Image processing device, image processing method, and program

    JP2022011704A

  • Computer Eye (PCEYE)

    JP2022172179A