Systems and methods for camera operation in a live environment
The system addresses latency in live camera tracking by using AI/ML to determine time-based differences between live and recorded musical scores, automating camera adjustments for precise and cost-effective live video production.
Patent Information
- Application Number
- JP2025513321
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-01
- Filing Date
- 2023-09-01
- Publication Date
- 2025-09-11
AI Technical Summary
Existing camera operation systems in live environments struggle with latency in tracking live audio and fail to efficiently adjust camera movements based on time-based differences between live and recorded musical scores, leading to costly and uncertain video production.
A system and method utilizing AI/ML to determine time-based differences between live and recorded musical scores, enabling automated camera selection and adjustment by inferring tags from these differences, allowing for precise camera movements and reducing the need for manual intervention.
Reduces costs and uncertainties in live video production by enabling reproducible camera adjustments and immersive experiences, even with varying live performances, by using AI/ML to align live and recorded audio frames for camera operation.
Smart Images

Figure 2025530121000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is a PCT application of U.S. non-provisional application 18 / 241,708, entitled "SYSTEM AND METHOD FOR CAMERA HANDLING IN LIVE ENVIRONMENTS," filed September 1, 2023, and is related to and claims the benefit of priority from U.S. provisional application 63 / 403,540, entitled "SYSTEM AND METHOD FOR CAMERA HANDLING IN LIVE ENVIRONMENTS," filed September 2, 2022, the disclosures of both of which are incorporated herein by reference in their entireties for all intents and purposes.
[0002] background 1. Field of the Invention
[0002] The disclosure herein relates generally to camera operation in a live environment, and more particularly to addressing latency in tracking a live version of audio in a live environment and the benefits of directing a camera based on a live version of audio in a live environment. [Background technology]
[0003] 2. Description of the Prior Art Audio recognition allows computer-based applications to determine the identity of a song by capturing a live version of the song. For known musical scores, computer-based applications can detect the artist or score using short recordings of live or recorded versions of the score from output devices such as radio or television, or even from a person. Certain computer-based applications can use humming or other vocalizations to attempt to detect the score. Summary of the Invention
[0004] A processor-implemented method includes determining, by at least one processor, that at least one frame of a live version of a musical score being performed in a live environment corresponds to a recorded frame of a recorded version of the musical score, the method further including enabling, by at least one camera, camera movement according to a time-based reference between the at least one frame and the recorded frame and according to at least one dominant influence of the recorded version or the live version.
[0005]
[0005] A system is provided for tracking a live version of a musical score to enable camera operation in a live environment. The system can be associated with at least one camera, a memory, and at least one processor for executing instructions stored in the memory to perform steps or functions, including determining that at least one frame of a live version of a musical score being performed in a live environment corresponds to a recorded frame of a recorded version of the musical score. The steps or functions further include enabling, by at least one camera, a camera movement according to a time-based reference between the at least one frame and the recorded frame and according to at least one dominant influence of the recorded version or the live version.
[0006]
[0006] Various embodiments according to the present disclosure are described with reference to the drawings. [Brief explanation of the drawings]
[0007] [Figure 1] Figure 1 shows an example of a live environment for camera operation. [Figure 2] FIG. 2 shows an example of a system for camera operation in a live environment. [Figure 3A]3A and 3B show example plots associated with methods and systems for camera operation in a live environment. [Figure 3B] 3A and 3B show example plots associated with methods and systems for camera operation in a live environment. [Figure 4A] FIG. 4A shows a view from camera operations in a live environment, where one or more cameras are selected or adjusted to point in a direction associated with at least one dominant influence. [Figure 4B] FIG. 4B illustrates aspects associated with an immersive experience in a live environment based in part on tracking a live version of the audio in the live environment. [Figure 5A] FIG. 5A illustrates an example of a method for camera operation in a live environment. [Figure 5B] FIG. 5B illustrates another example of a method for camera operation in a live environment. [Figure 6] FIG. 6 shows a further example of a system for camera operation in a live environment. DETAILED DESCRIPTION OF THE INVENTION
[0008]
[0015] Various embodiments are described below. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that embodiments may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments. Various other features may be implemented in various embodiments and are discussed and suggested elsewhere herein.
[0009]
[0016] Additionally, the disclosure herein generally relates to cameras and devices with cameras, including smartphones, standalone cameras, laptops, notebooks, and the like. Thus, reference to one or more cameras also includes such other devices. The disclosure herein addresses the operation of such cameras in a live environment, such as during a live classical music performance of a piece or composition. However, the disclosure herein is also applicable to concerts of other musical genres with multiple performances and cameras capturing various views of the concert. By way of example, camera operation herein addresses the time-based difference between a live version of audio (e.g., corresponding to a particular live classical music piece or performance) and a recorded version (also referred to herein as a reference audio recording) in real time.
[0010]
[0017] Recorded versions may have many differences, such as different interpretations by various conductors, and may be from studio recordings or different performances of the same particular classical music piece. Differences from different interpretations may also include different lengths (such as different total audio lengths) of the same particular classical music piece. Certain approaches may enable audio recognition using computer-based applications to determine songs by capturing live versions of songs, but these do not involve determining time-based differences to enable camera operations. For example, computer-based applications may enable the detection of artists, musical scores, and other relevant information using recorded versions or even portions of live versions in live environments, such as from live performances or output devices including radio and television, or from human vocalizations. Certain computer-based applications may enable the recognition of humming or other human vocalizations for use in musical score detection, for example. In contrast, the approach herein involves using AI / ML to determine the time-based differences between a recorded version of a musical score (such as a particular classical music piece) and a live version of the same piece, where, for example, the piece may be interpreted differently by a conductor, soloist, or orchestra, and the piece may consequently vary in performance length.
[0011]
[0018] The system and method described herein addresses deficiencies raised and mentioned elsewhere herein by providing a system and method for camera operation in a live environment. Video production in a live environment (such as a classical music performance or concert) can be a costly endeavor, requiring a team of trained personnel and resulting in uncertainties. The method or system described herein allows a user to pre-stage a concert, such as by using a recorded version of a musical score with associated tags and performed in the live environment, tracking time-based differences, and using the time-based differences and at least one dominant influence from the tags of the live version to handle camera selection (from multiple different cameras) or camera adjustment (changes in pitch, tilt, or zoom) of one or more cameras in the live environment.
[0012]
[0019] Thus, a method or system herein is for an automated approach to enable camera selection or camera adjustment of one or more cameras during a live performance, such as during a concert that is representative of a live environment. This approach allows pre-defined directions in tags associated with a recorded version to be modified for use with one or more cameras. For example, the pre-defined directions may be associated with a recorded version of a musical score. Furthermore, the pre-defined directions may be associated with at least one dominant influence in the musical score, such as which section of the orchestra (wind section, string section, or percussion section) or which musical feature (rhythm, tone, pitch, timbre, intensity, etc.) is dominant in the recorded version and at which reference moment in the score. Thus, the pre-defined directions can be used to direct one or more cameras during a live version, based in part on input regarding where the live version is in the score relative to the recorded version.
[0013]
[0020] Also, the methods or systems herein may allow for initial camera selection and camera adjustment using at least one dominant influence from the recorded version of the live performance, for example. Alternatively, in this approach, the recorded version of the score may be used to ensure that any time-based differences relative to the live version can be addressed by overriding the initial camera selection and camera adjustment with further camera selection and camera adjustment from the live version of the live performance, or by overriding the initial camera selection and camera adjustment with default production settings for camera selection and camera adjustment. For example, the default production settings for camera selection and camera adjustment ensure that the conductor's view is provided when there are unknown time differences, such as delays or extensions in the music, in the live version compared to the recorded version of the score.
[0014]
[0021] The method or system herein can reduce costs by eliminating intervention by camera operators, on-site music readers, and art directors for each of one or more cameras, and can also enable reproducible results in the displayed video of various live performances of a musical score, even if changes occur in the live version. Furthermore, while musicians are positioned on stage according to a stage plan, the use of multiple cameras allows musicians, or even sections of musicians, to be positioned so that the display of the live version of the musical performance coincides with the stage plan. The camera direction herein allows for following artists or sections as they move around the stage.
[0015]
[0022] The methods or systems herein include one or more cameras that can be pointed in a direction associated with at least one dominant influence, such as a flute solo in a wind section, thereby dominating the flutist and other such individual musicians or sections when necessary. Furthermore, the approaches herein not only consider frequency and rhythm to determine at least one dominant influence, but can also use timbre in making such determinations. The methods or systems herein also incorporate speed improvements for video tracking of live performances for display, for example, by selecting the smallest time-based difference between multiple recorded versions of the same musical score played in a live environment.
[0016]
[0023] The method or system herein includes an artificial intelligence / machine learning engine (AI / ML engine) trained to use time-based differences determined between the live and recorded versions of a musical score to infer tags that are sent to an automated production module. The AI / ML engine herein converts data associated with the live and recorded versions of the musical score into numerical vectors. For example, tokenization can be performed to separate the data into multiple portions, such as one or more frames of the musical score. These portions can then be further subdivided. The method or system herein enables tokenization to generate tokens, which are communicated to the automated production module to enable camera selection and camera adjustment, thereby generalizing the relationship between the data and the tags from the recorded version associated with the data. As a result of the tokenization, a set of unique tags can be generated for the data and associated with the tokens.
[0017]
[0024] Additionally, vectorization can be performed as part of the AI / ML engine. Vectorization of tokens associated with time-based differences between the live and recorded versions provides a numerical vector. The numerical vector can be used in feature selection to determine which features to retain for use in automatically assigning tags to time-based differences between the live and recorded versions of the score. The AI / ML engine herein includes at least temporarily storing distances between tokens. Such distances can be used to reduce the number of vectors required and can also reduce the need for vector comparisons as part of the AI / ML engine to assign tags to time-based differences.
[0018]
[0025] In at least one embodiment, a separate AI / ML engine can be used in a similar fashion to extract features for use in the frame referencing module, providing time-based differences between versions of a musical score. Again, the AI / ML engine for frame referencing herein includes at least temporarily storing distances between tokens from the tokenization process. Such distances can be used to reduce the number of vectors needed to extract features and can reduce the need for vector comparisons as part of the AI / ML engine to determine features for use in the frame referencing module.
[0019]
[0026] Furthermore, a live version of a score can be analyzed simultaneously with respect to multiple recorded versions using such an approach. This can be done using parallel processing to increase robustness and tolerance to various interpretations of the score in a live environment. The method or system herein may only require comparison of an individual vector of one dominant influence (e.g., tone) with the matrix of the same dominant influence from multiple recorded versions, each referenced by a reference number in the matrix. Furthermore, pooling can be used to reduce the size of the matrix.
[0020]
[0027] A system associated with the methods herein includes a memory and at least one processor for executing instructions stored in the memory to perform the method steps. The method steps include analyzing audio from a live version (e.g., live audio from a live interpretation of a classical music piece) with a recorded version (e.g., pre-recorded audio from a different interpretation of the same piece). The analysis is to determine time-based differences that can be expressed over a time interval. In at least one embodiment, the time-based differences are from at least one frame of each live version (live interpretation) relative to the recorded version (pre-recorded interpretation). Furthermore, the time-based differences can be used in conjunction with at least one dominant influence (e.g., a flutist playing a solo in a classical music piece) to enable camera selection or adjustment of one or more cameras during a live classical music performance. In this approach, one or more cameras can be directed toward at least one dominant effect in the live environment based on at least one dominant effect associated with the recorded version, even if there is a time-based difference in where the at least one dominant effect occurs between the recorded and live versions of the same classical music piece.
[0021]
[0028] FIG. 1 illustrates an example live environment 100 for camera operations, which are the subject of the present description. While the live environment 100 is illustrated as a classical arrangement for an orchestra, live concerts of any genre can benefit from the present disclosure. The illustrated live environment 100 includes a string section 102, 104, 106 having at least violins and cellos; a woodwind (or brass) section 114 having, for example, a clarinet, bass clarinet, and bassoon; a brass section 118 having tubas and trumpets, which may in some cases be grouped with the brass instruments; a percussion (or bass) section 108; a vocal section 110; an auxiliary section 116; and a conductor's podium 112. The string sections 102, 104, 106 may be the largest sections in a classical arrangement for an orchestra. Additionally, the auxiliary section may include piano, harp, and other large, specialized instruments. The orchestra is on stage 130, which may be an entirely live environment.
[0022]
[0029] The methods or systems herein allow a conductor to use each section differently in various parts of a score. Some scores may require multiple string section personnel in a single part, including a specific string section, such as violins or cellos, or a combination. Furthermore, some scores may include a combination of multiple sections and individual performances of each section. The methods or systems herein may provide multiple cameras 122-130, which may be directed at a single section, a single person, or may be enabled for specific sections of specific people in a live environment. The cameras may capture the live environment for display in real time or record the live environment for later viewing. Furthermore, one or more cameras 122-130 may be directed toward the audience or spectators 120.
[0023]
[0030] FIG. 2 illustrates an example system 200 for camera operation in a live environment. The system 200 includes at least one processor and a memory having instructions that execute on the at least one processor to perform the functions described herein. Further description of such an example system may be apparent from a review of FIG. 6 herein. The system includes a live audio / video feedback line 214 for providing a live version of the musical score from the live environment through a live version module 202 to a sound tracking module 208. The live version module 202 includes at least a memory for providing the live version in real time to the sound tracking module 208.
[0024]
[0031] The system includes a recorded versions module 204 for providing recorded versions of a musical score to the sound tracking module 208. Accordingly, the recorded versions module 204 includes a library of recorded versions of a musical score and predetermined information 222 for each recorded version. The predetermined information 222 may include tags, such as metatags for dominant influences at various points or epochs within the recorded version of the musical score. Furthermore, the recorded versions module 204 may also maintain live versions of various interpretations of the musical score, which would result in different recorded versions of the same musical score, providing further robustness to the system 200 herein.
[0025]
[0032] Thus, predetermined information 222 can be created and stored with the various recorded versions. In application, various recorded versions of the same musical score can be used to select one recorded version with the smallest time-based difference from the live version using the approach herein. For example, parallel processing can be used to process at least one frame of each different recorded version in real time in parallel with frames of the live version. Thus, a reliable recorded version (reference audio) can be selected from the various recorded versions based at least in part on a reliability factor. In one example, the time-based difference can be generated and minimized to a minimum value, where the minimum value represents the reliability factor. Thus, one recorded version with the smallest time-based difference can be selected so that its tag can be used with the live version. Thus, the systems and methods herein can respond more quickly to handle camera selection and adjustment according to the smallest time-based difference. In the methods or systems herein, the live version and the recorded version can be provided as digital data in a combined form to the vectorization module 206A.
[0026]
[0033] The sound tracking module 208 can vectorize the digital data via a vectorization module 206A implemented by at least one processor. The vectorization module 206A can also be used to generate vectors based on tokens for use in the frame reference AI / ML engine 206B, which can also be used in the tagging-related AI / ML engine 210. The tagging-related AI / ML engine 210 can separately tokenize and vectorize the data primarily using the time-based differences 216 determined by the sound tracking module 208. The frame reference AI / ML engine 206B generates features from one or more of a recorded or live version of the musical score. The features are provided to the frame reference module 206C.
[0027]
[0034] The frame referencing module 206C of the sound tracking module 208 performs frame referencing using a suitable algorithm, such as using online dynamic time warping (ODTW). The frame referencing module 206C provides time-based differences 216 from the output of the sound tracking module 208 to the tagging-related AI / ML engine 210, which then infers or predicts tags 218 to be provided to the live version in real time. The tagging-related AI / ML engine 210 can learn how to tag the live version based in part on the time-based differences 216 between the recorded and live versions of the musical score and in part on tags from the predetermined information 222. The tagging-related AI / ML engine 210 can then associate those tags from the predetermined information 222 of the recorded version with frames of the live version according to the time-based differences 216.
[0028]
[0035] The tags 218 can be provided to the auto-staging module 220 to identify at least one dominant effect (e.g., a flute solo) in the recorded version of the score that is also present in the live version of the score. Taking into account time-based differences between the live and recorded versions of the score, the tagging-related AI / ML engine 210 can inform the auto-staging module 220 of the tags 218. For example, an initial time-based difference (e.g., a two-second delay) can be learned after the live version begins, and the tagging-related AI / ML engine 210 can project the tags 218 from predetermined information 222 onto the rest of the live version, thereby appropriately directing the cameras 122-130 at future times and instantaneously.
[0029]
[0036] Additionally, the tagging-related AI / ML engine 210 can beneficially take into account ongoing time-based differences as the live version continues. Thus, the tagging-related AI / ML engine 210 can be trained using the time-based differences to infer or predict further points in the live version that correspond to the recorded version, thereby signaling tags to the auto-staging module 220 in anticipation of those upcoming points in the live version. As the live version may have further time-based differences, the tagging-related AI / ML engine 210 can update the tags 218 for those points as the time-based differences change. The tagging-related AI / ML engine 210 can also be used to create tags based on provided predetermined information 222. These tags 218 can be, for example, a combination of one or more tags from the predetermined information 222. The tags 218 can be converted into specific camera presets in the auto-staging module 220, which can send signals to one or more cameras for selection or adjustment as described throughout this specification. In one example, the NN of the tagging-related AI / ML engine 210 receives predetermined information 222, which may include a stage plan, and metadata (with tags) that are predetermined and associated with the recorded version.
[0030]
[0037] Thus, once trained, the tagging-related AI / ML engine 210 can infer tags 218 associated with future points in one or more frames of the live version of the musical score. The tagging-related AI / ML engine 210 communicates the tags 218 to the auto-staging module 220. The tags enable the auto-staging module 220 to send signals to one or more cameras 122-130. The signals enable camera selection or camera adjustment of one or more cameras 122-130 in the live environment. The live version of the musical score can be analyzed frame-by-frame by the sound tracking module 208 and the tagging-related AI / ML engine 210. Time-based differences 216 are communicated to the tagging-related AI / ML engine 210 to analyze each received frame of the live version against one or more frames of one or more recorded versions of the musical score in the live environment.
[0031]
[0038] The recorded version module 204 includes predetermined information 222, which indicates at least one dominant influence in the musical score. Thus, once a time-based difference 216 (delay, slow tempo, etc.) is determined, this predetermined information 222 can be passed to the tagging-related AI / ML engine 210. The tagging-related AI / ML engine 210 uses this predetermined information 222, along with the time-based difference 216 provided by the frame referencing module 206C, to determine tags 218 that will be communicated to the auto-staging module 220. The tags 218 communicated to the auto-staging module 220 enable the module 220 to use the time-based difference from the frame-by-frame analysis to select or adjust one or more cameras 122-130 in the live environment to point in a direction associated with at least one dominant influence. Thus, staging can be instantaneous or anticipated, since the tags 218 can include future tags for a predetermined time beyond the current point in the live version. Thus, the system 200 can determine delays or changes in the rhythm of the live version of the musical score, can determine at least one dominant effect that is slowed, sped up, or otherwise different, and can select or adjust at least one camera to address the at least one dominant effect.
[0032]
[0039] The methods or systems herein include feature extraction for the live or recorded version, which may be performed in a frame-reference AI / ML engine 206B separate from the tagging-related AI / ML engine 210 associated with the tagging approach described herein. The sound tracking module 208 may include the frame-reference AI / ML engine 206B for feature extraction. For example, the frame-reference AI / ML engine 206B may include an autoencoder to perform such feature extraction. For example, the autoencoder may be a vector quantization variational autoencoder (VQ-VAE). The VQ-VAE attempts to find a latent space (e.g., variables) that have an underlying representation in the provided data, such as data associated with the live or recorded version. The VQ-VAE maps the live and recorded versions from the received digital data in a frame-by-frame manner, performs tokenization as described elsewhere herein, and provides vectors associated with the tokens. The tokenization may be represented by embedding indices. The approach here is to use VQ-VAE together with ODTW to select features and to use those features to determine time-based differences.
[0033]
[0040] The frame-reference AI / ML engine 206B can have stored embeddings, thereby eliminating the need for vector comparisons to learn features used in the ODTW algorithm. Separately, a similar approach can be taken with the tagging-related AI / ML engine 210, which can learn how to tag the live version based on time-based differences and tags from the recorded version and provide the tags to be passed to the auto-staging module 220. The frame-reference AI / ML engine 206B and the tagging-related AI / ML engine 210 can be periodically updated as an average of recently tokenized data. The method or system herein can use a moving average in one approach. Furthermore, the frame-reference AI / ML engine 206B can be trained to estimate differences between the live and recorded versions of a musical score. The tagging-related AI / ML engine 210 is used to infer tags that are passed to the auto-staging module 220.
[0034]
[0041] The neural network (NN) of the tagging-related AI / ML engine 210 can learn from the time-based differences of the sound tracking module 208 to dynamically align the input layer to weights for the described tags available within the NN. Furthermore, the frame-reference AI / ML engine 206B can also dynamically update its input layer to dynamically align weights for the described features of the audio version within its NN. Tags can be defined for use with various dominant influences and can be generated by an adversarial neural network. For example, a discriminative network of such an adversarial neural network can be trained to classify forced variants of the recorded version, and a generative network can be used to generate the forced variants. The forced variants can have specific dominant influences of known tags from the predefined information 222, which are informed post-training of the discriminative network. Thus, in this approach, specific dominant influences are automatically tagged, and tags can be generated for use by the NN of the tagging-related AI / ML engine 210.
[0035]
[0042] The weights for the NN of the tagging-related AI / ML engine 210 can be updated according to time-based differences and used by neurons in subsequent layers. For example, if the ODTW analysis indicates that there are no differences between the versions, the weights of the NN will not be changed, and no tags will be assigned or reported by the NN as frames of the live version of the musical score are processed by the AI / ML engine 210. This can be a feed-forward approach for the NN of the tagging-related AI / ML engine 210. With respect to differences identified by the ODTW analysis, the NN can be skewed so that its output identifies at least one tag associated with the difference and can be associated with the dominance of the live or recorded version. Thus, time-based differences in the versions can be automatically handled and tagged without having to address the specific dominance of the live version relative to the recorded version.
[0036]
[0043] 2 , the system 200 can include at least one processor as part of the auto-staging module 220, the tagging-related AI / ML engine 210, and / or the immersive experience module 224. The at least one processor can make a determination that at least one frame of a live version 202 of a musical score being performed in a live environment 212 corresponds to a recorded frame of a recorded version 204 of the musical score. At least one camera 122-130 can then be enabled to perform a camera movement according to a time-based criterion, including a time-based difference 216 between the at least one frame and the recorded frame, and according to at least one dominant influence of the recorded version 204 or the live version 202.
[0037]
[0044] 2, the dominant influence may be in the form of instructions stored in a memory associated with the at least one processor or in a separate memory, so that the recorded version 204 of the musical score may be accessed from such memory. The instructions may be executed by the at least one processor. Furthermore, the instructions related to the dominant influence may be different from the instructions for the functions performed by the at least one processor to determine the at least one frame and to perform the camera movement. The instructions associated with the dominant influence may include at least a first instruction related to at least one of the musical performers to be captured by the at least one camera 122-130 using the camera movement for the at least one frame, or may include at least a second instruction related to the type of immersive experience to be generated and captured by the at least one camera 122-130. Additionally, further steps or functions performed by at least one processor, such as those of the auto-staging module 220, include enabling camera movements to capture musical performers or immersive experiences, and incorporating or including camera movements in the live version 202 of the live environment 212.
[0038]
[0045] In a further example, using system 200 in FIG. 2, instructions regarding the dominant influence can be provided in a computer-based input. For example, a character or graphical user interface can be used to receive the computer-based input using at least one component of system 600 in FIG. 6. The dominant influence provided in the computer-based input can be based on the recorded version 204 of the musical score. In one example, the immersive experience is generated for at least one frame of the live version 202. Furthermore, the immersive experience can be of some type, such as at least one of the graphical visualization 452 in FIG. 4B displayed during the live performance and the text 454 displayed to the audience of the live performance. For example, autostaging module 220 provides input to immersive experience module 224 to cause the projection 226 or other form of immersive experience feature to provide the graphical visualization 452 or the text 454. Although a particular section view 410 (of the percussion section 410) is illustrated, the immersive experience can be provided in any part of the live environment 212 / 402 itself.
[0039]
[0046] The system 200 also supports adjusting camera movement or the immersive experience in response to features obtained from the audio and video of the live performance audio / video feedback line 214. For example, such features may include at least one of the volume or level of emotional expression of the musical performer. Adjusting may include at least one of zooming a camera associated with at least one of the cameras 122-130 and tuning the color of the graphical visualization 452 or text 454. In one case, the level of emotional expression of the musical performer may be determined at least in part based on an object tracking AI / ML algorithm in the object tracking module 228 that tracks 230 the musical performer's facial matrix in the live version of the musical score. This AI / ML algorithm may be distinct from the frame-reference AI / ML engine 206B and the tagging-related AI / ML engine 210. The object tracking AI / ML algorithm may be trained using facial matrices for different emotional expressions. The object tracking AI / ML algorithm can determine various emotional expressions associated with a musical performer by receiving a filtered (or processed in any manner that extracts facial details) facial matrix from a capture of the musical performer and classifying the facial matrix into at least one of various emotional expressions.
[0040]
[0047] In one aspect, adjusting the camera movement or immersive experience is further based in part on object tracking 230 performed on the live version 202 of the musical score. For example, the live version 202 may be the input to the object tracking module 228, allowing for fast, near-real-time processing. This differs from a live version captured in real time and may be the result of one or more additional video processing steps already performed on the live version in the audio / video feedback line 214. For example, the object tracking module 228 may need to perform additional processing when directly capturing the live environment 212, as illustrated, for example, by the dashed lines to the object tracking module 228.
[0041]
[0048] Additionally, the system 200 may include enabling camera movement using multiple of the at least one camera 126, 128. For example, the camera movement may include specific actions performed by specific ones of the at least one camera 122, 124, 126, 128, 130, such as to provide specific section views 406-414 including views of the string section 408, the brass section 414, and the percussion section 410. The camera movement may be altered in real time based at least in part on a determination of a problem with at least one camera or additional cameras associated with that at least one camera. For example, if one camera is not operational, the autostaging module 220 may use the remaining ones to compensate for the loss of one of the specific section views 406-414.
[0042]
[0049] The system 200 also supports determining, with respect to at least one frame of the live version of the musical score, a recorded frame of the recorded version of the musical score based on the frame-referenced AI / ML engine 206B, which can be used to extract features describing at least one frame of the live version 202 and the recorded frame of the recorded version 204. The frame-referenced AI / ML engine 206B can use online dynamic time warping (ODTW) to determine associations between audio features.
[0043]
[0050] Here, one or more of the AI / ML engines 210 may be based at least in part on feedback provided by a user. For example, the feedback may be collected via computer-based input, including errors in frame matching. A user may provide such computer-based input remotely, for example, using a handheld device as described in connection with FIG. 6 . The feedback may be added to a training data set for the AI / ML engine 210. The systems herein enable any of the AI / ML algorithms herein, such as those in example modules 206B, 210, and 228 of FIG. 2, to use results from a previous live performance of the musical score. For example, using at least one frame to determine the recorded frame allows consecutive frames to be matched between the live and recorded versions.
[0044]
[0051] Any of the AI / ML algorithms herein may use a confidence level, where a first predetermined value of the confidence level allows the AI / ML algorithm to influence the camera movement by one or more of switching to another dominant influence type or setting and maintaining at least one camera 122-130 or multiple cameras in a switched mode until the machine learning algorithm regains a second predetermined value of the confidence level for the next frame of the live or recorded version of the musical score, where switching may include at least one of directing at least one of the at least one camera 122, 124, 126, 128, 130 or two or more cameras 122-130 to capture the conductor of the musical score based on an immersive experience associated with the live version of the musical performance and reducing an influence associated with at least one dominant influence.
[0045]
[0052] Any of the AI / ML algorithms herein can take the form of an ensemble of models, with each model in the ensemble taking as input the same or a different set of features extracted from the audio signal on the audio / video feedback line 214. Each model in the ensemble can apply ODTW individually. The individual outputs of the ODTW procedures for each model in the ensemble can be combined together, and a confidence score can be calculated for the combined output. The combined output of two or more models in the ensemble can be an average of the outputs. In one example, the confidence score can be based in part on comparing whether the outputs of two or more models in the ensemble are different from each other. Furthermore, if all or nearly all of the outputs are the same (or similar, for example, according to a predetermined variable), the confidence score for two or more models in the ensemble can be equal to or nearly equal to a maximum value, such as the numerical value "1."
[0046]
[0053] In a further example, the more the outputs of two or more in an ensemble of models differ from each other, the more likely they are to be determined with a lower confidence level. Furthermore, if the confidence level is too low with respect to a predetermined value (also referred to herein as a confidence tolerance threshold), which may be associated with manual or automatic tuning, the ensemble of models may report a problem with certainty about its output. This allows the system 200 to operate in an error scenario, taking safety actions such as providing a default production setting that captures the view of the conductor, audience, or entire orchestra instead of the individual musicians. This safety action may be appropriate until the ensemble of models begins to deliver outputs with a confidence level that exceeds the predetermined value.
[0047]
[0054] The recorded versions 204, one of which is illustrated in FIG. 2, may be various versions of a musical score, which may be stored in system 200 (and / or system 600 of FIG. 6) along with their associated dominant influences. A recorded version may be determined from multiple recorded versions to be used with the live version at the beginning of the musical score. This determination may be based in part on differences, such as time-based differences 216, which are part of a time-based measure between at least one frame and each of the recorded frames of the recorded version.
[0048]
[0055] In some examples, a recorded version in the plurality of recorded versions may be determined to be used with the live version of the musical score based in part on a confidence level that causes an AI / ML algorithm, such as one or more of the frame reference AI / ML engine 206B or the tagging-related AI / ML engine 210, to provide output using at least one frame for the plurality of recorded frames. For example, one or more of the frame reference AI / ML engine 206B or the tagging-related AI / ML engine 210 may be responsible for determining the recorded frame associated with at least one frame and for making such a determination by providing a confidence level. The confidence level may be obtained along with the recorded frame determination to determine whether to keep or change the recorded frame. Even if a recorded frame determination is made, the provided confidence level may be lower than a predetermined value. Thus, the recorded frame determination to be used may be rejected in favor of another recorded frame of a different recorded version having a higher confidence level relative to the predetermined value, which is then used with the live version.
[0049]
[0056] The system 200 includes using an auto-staging module 220 to enable camera movement of at least one camera or cameras 122-130 using preliminary instructions until a recorded version is determined with a predetermined value of confidence. The determination of the recorded version can be performed in real time or near real time using a scalable cloud environment, such as the system 600 of FIG. 6, and with multiple recorded versions available to the system 600. Accordingly, information associated with the recorded version can be sent to at least one camera or multiple cameras 122-130. The system 200 further supports enabling the further real-time determination of another recorded frame to be used with the live version, by one or more of the frame reference AI / ML engine 206B and the tagging-related AI / ML engine 210, using a scalable cloud environment, such as the system 600 of FIG. 6, and an associated edge environment.
[0050]
[0057] Additionally, system 200 supports capturing a live version using at least one camera or multiple cameras 122-130 in real-time synchronization with live audio, such as via provided feedback feature 214. Synchronized capture can be used in at least one of online broadcasting, television-based live entertainment production, and production of sheet music recordings. Thus, camera movement is enabled to obtain various live versions of a sheet music for various playbacks, such as for online broadcasting, television-based live entertainment production, and production of sheet music recordings.
[0051]
[0058] 3A and 3B show example plots 300, 350 associated with camera operation in a live environment according to at least one embodiment herein. A method or system herein includes a dynamic time warping (ODTW) algorithm for nonlinearly normalizing the time between at least two frames of multiple versions of a musical score. Furthermore, the ODTW works optimally, for example, to provide at least a sum of distance values for sequentially aligned or ordered index pairs. The ODTW provides statistical measures that benefit from optimization for alignment or ordering. This can be part of the dynamic programming performed by the ODTW. For example, the ODTW compares a numerical representation of the live version with a stored recorded version.
[0052]
[0059] In this example, the recorded and live versions of a musical score are received as a numerical representation of the score. The numerical representation of the versions may rely on specific coefficients, such as chroma features, cepstral coefficients, Mel-frequency cepstral coefficients, and perceptual linear prediction (PLP) coefficients. Furthermore, feature extraction via a frame-reference AI / ML engine uses pattern comparison by searching for expressions between versions of a musical score. FIG. 3A illustrates a plot 300 of frame coefficients of a recorded version 308 versus a live version 310 over time 306, which represents a reference epoch of the musical score. Furthermore, the plot 300 shows y-axes 302, 304 as having normalized values of the musical score relative to the recorded version 308 and the live version 310, including at least one influence that can be used, for example, to determine a dominant influence in the recorded version. Furthermore, a mapping 312 is shown between the versions 308, 310, representing time-based differences between the versions.
[0053]
[0060] The methods or systems herein use ODTW to optimize the distance between various coefficients at one or more time points 306. For example, the approach herein optimizes the distance between coefficients (points on each line 308, 310) by copying and mapping each such point for one or more time points 306. Additionally, a sliding window can be defined and used with each point of the live version 308 for all points within that sliding window for the recorded version 310. The cost between such points is determined using a cost function of the ODTW.
[0054]
[0061] Furthermore, minimization of the cost function can be used to determine the distance (and minimum distance) between the current point and the initial point in the sliding window. This is done similarly for the initial point to the matching point in the recorded version 308. A cost function can be obtained that has a cost value for the final point in the sliding window, and therefore the minimum distance between versions can be determined using the cost value obtained from matching the final points between the versions in the sliding window until the end of the frame in a frame-by-frame analysis. Because the recorded version is known relative to the live version, the estimated start time can be used to determine the distance between points on the lines of versions 308, 310. With the time information relative to the x-axis used to determine the time-based difference, the distance can be minimized at each frame to determine where the points in the live version are estimated to be relative to the recorded version. Furthermore, because the time-based difference carries over throughout the score, one frame can be used in subsequent frame analyses.
[0055]
[0062] In the example plot 350 associated with camera operation in a live environment, minimized differences and time-based differences between versions of a musical score are provided for illustrative purposes. The methods or systems herein process numerical representations without generating such plots. For example, the x and y axes 352, 354 of plot 350 relate to ranges for sequences of coefficients used in the ODTW analysis of versions of a musical score. These ranges may be approximately the same. In plot 352, proportional linear sections 356 imply agreement between versions, and any skew 258 represents a contraction or expansion of the time-based differences between versions with respect to at least one frame.
[0056]
[0063] FIG. 4A illustrates an example camera operation 400 in a live environment in which, according to at least one embodiment herein, one or more cameras are selected or adjusted to point in a direction associated with at least one dominant influence in the live environment. With respect to the live environment 402, at least one camera provides a default view 404 and multiple specific section views 406-414. The specific section views 406-414 include views of a string section 408, a wind section 414, and a percussion section 410. Furthermore, if a single camera is used for two of the views 404, 406-414, that single camera may be subject to adjustment using time-based differences and at least one dominant influence. However, multiple cameras may also be selected and adjusted using time-based differences and at least one dominant influence. Furthermore, a single camera may be selected based on a desire to make adjustments, such as pan, tilt, zoom, etc., to focus on one or more objects in the live environment 402.
[0057]
[0064] The methods or systems herein may, for example, allow for initial camera selection and camera adjustments using at least one dominant influence in a recorded version of a live performance. These initial camera selections and camera adjustments may illustratively be for providing a view to an audience on a display screen or for recording along with the live version of the score. The recorded version of the score may be used to ensure that any time-based differences relative to the live version are addressed by overriding the initial camera selections and camera adjustments with further camera selections and camera adjustments from the live version of the live performance, or by overriding the initial camera selections and camera adjustments with default direction settings for camera selection or camera adjustment, such as default view 404. For example, the default direction settings for camera selection or camera adjustments may ensure that a conductor's view is provided when there are unknown time differences, such as delays or extensions in the music, in the live version compared to the recorded version of the score.
[0058]
[0065] 4B illustrates an aspect 450 associated with an immersive experience 452, 454 in a live environment based in part on tracking a live version of audio in the live environment. For example, a camera can be automatically aimed at an object 456 (such as a particular musical performer) to provide the immersive experience 452, 454 in real time using the provided immersive experience module 224. The musical performer 456 can be envisioned moving over time relative to the live version being played. In one example useful for the immersive experience 452, 454 herein, the system 200 can use its object tracking module 228 and cameras 122-130 to track the musical performer 456 to provide the immersive experience 452, 454. However, if the object tracking module 228 loses track of the musical performer 456, the system 200 can switch to a different operating scenario associated with camera operation as a backup, while still maintaining the ongoing immersive experience 452, 454 using the immersive experience module 224.
[0059]
[0066] In another example, the object tracking module 228 can include emotion detection related to the emotional expression of the musical performer 456. The output associated with the emotional expression can be used to provide additional dominant influences that affect camera operation. For example, zooming in or out on at least one of the provided cameras 122-130 can be based on the detected emotional expression of the musical performer 456. Furthermore, switching camera operation to a default performance setting can occur when the emotional expression of the musical performer 456 is negative.
[0060]
[0067] System 200 can also be integrated with software for various types of presentation or display media. For example, immersive experience module 224 can provide graphical visualization 452 and text 454 (reflecting words, lyrics, names, etc., such as "THE BAND"), but can also include any other software used to generate an immersive experience other than graphical visualization 452 and text 454. Furthermore, graphical visualization 452 can move with musical performers 456 and / or live versions of audio and can include three-dimensional (3D) footage, faces, images, cartoons, and other visualizations. Thus, the graphical visualization, text, or any other form of displayed media of the immersive experience herein can be synchronized with one or more of live audio and live video from the live environment. The displayed media can vary depending on features derived from the live audio. For example, during a live performance, when a particular dominant effect occurs, graphical visualization software of the immersive experience module 224, which is connected to a display screen in the live environment 212, can project a corresponding graphical visualization onto a display screen used with the system 200 or through an augmented or virtual reality device.
[0061]
[0068] FIG. 5A illustrates an example method 500 for camera operation in a live environment. Method 500 includes determining (502) a musical score to be performed. Once performance begins, method 500 includes analyzing (504) the live version of the musical score with a recorded version of the musical score in the live environment. Analyzing (504) may include the On-Demand Two (ODTW) approach described throughout this specification. For example, analyzing (504) may include determining, by at least one processor, that at least one frame of the live version of the musical score being performed in the live environment corresponds to a recorded frame of the recorded version of the musical score. A time-based difference is determined (506) from the analysis in step 504. Further, verification (508) of at least one dominant influence in the live version or the recorded version may be performed.
[0062]
[0069] A dominant influence can be predetermined and associated with the recorded version. For example, a solo by an orchestra member, a particular section's performance, and other such aspects can be a dominant influence that requires the camera to focus on that member or a particular section. However, when a dominant influence occurs with respect to the live version of the score, the time-based difference between the live and recorded versions of the score can change. The time-based difference forms the basis for the AI / ML engine to provide tags for enabling (510) camera selection or camera adjustment of one or more cameras in the live environment. For example, at least one processor uses the tags to determine one or more cameras to be selected or adjustments to be made to the one or more cameras. For example, the enabling (510) step can be performed by at least one processor and / or by at least one camera. The enabling (510) step can be for camera movement that is based on a time-based criterion between at least one frame and the recorded frame and that is based on at least one dominant influence of the recorded or live version.
[0063]
[0070] FIG. 5A further illustrates an example method 500 for camera operation in a live environment. Method 500 includes determining (502) a musical score (e.g., a musical composition) to be performed. Once the performance begins, method 500 includes analyzing (504) the musical composition as it is performed live in the live environment against a recorded version (e.g., a reference musical composition). Analyzing (504) may include the On-Demand Two (ODTW) approach described elsewhere herein. Time-based differences are determined (506) from the analysis in step 504. Furthermore, a dominant influence in the live version may be in the form of a delay or other change in the live version relative to the recorded version. The approach herein provides a default view for such a dominant influence in the live version of the musical score, such as by selecting a camera pointing toward the conductor and zooming, panning, or tilting the camera toward the conductor.
[0064]
[0071] FIG. 5B illustrates another example method 550 for camera manipulation in a live environment. Because at least the method of FIG. 5B changes the camera, method 500 may be an alternative to method 550 of FIG. 5A. For example, method 550 of FIG. 5B includes determining (552) a musical score to be performed. Once performance begins, method 550 includes determining (554) at least one dominant influence in the live or recorded version of the musical score. Further, method 550 includes analyzing (556) the live version of the musical score with the recorded version of the musical score in the live environment on a frame-by-frame basis. Analyzing (556) may include the On-Demand Two-Time approach described elsewhere herein. Verification (558) can be performed on time-based differences determined from the analysis in step 556. The method 500 further includes selecting or adjusting (560) one or more cameras in the live environment to point in a direction associated with at least one dominant influence using the time-based differences from the frame-by-frame analysis of step 556. The system 200, 600 herein supports turning on or off at least one function associated with the processor-implemented method 500, 550 using computer-based input reporting at least one aspect of the musical score. For example, at least one function that may be turned on or off is controlling the level of automation of camera movement.
[0065]
[0072] FIG. 6 illustrates another example system 600 for camera operation in a live environment. The system 600 may include computer and network aspects to perform at least some of the features described in FIGS. 2, 3, 5A, and 5B. The computer and network aspects 600 may include a distributed system. The distributed system 600 may include one or more computing devices 612, 614. The one or more computing devices 612, 614 may be adapted to run and function using a client application, such as a browser or a standalone application, and may be adapted to run and function over one or more networks 606.
[0066]
[0073] In at least one embodiment, a computing device herein can be a unit of a system for performing some of the features described herein, such as vectorization. However, such a computing device can also include a smartphone for performing at least some of the features described herein. For example, a smartphone with a camera can be used in place of a camera to receive instructions, such as adjustments or instructions for pointing the smartphone. In one example, there can be various types of devices that can be selected or adjusted for the camera operations described throughout herein, including devices for recording after being registered for use with the system herein. Furthermore, because certain live environments may involve movements involving one or more artists, it can be beneficial to incorporate other audience members into the system so that their cameras are integrated with instructions sent to their smartphones. The instructions at least instruct the adjustments that allow their smartphones to be used for the performance in a live environment.
[0067]
[0074] Additionally, the server 604 having components 604A-N can be communicatively coupled to computing devices 612, 614 via a network 606 and, if provided, a communications device 608. The components 612, 614 include a processor, memory, and random access memory (RAM). The server 604 can be adapted to perform services or applications associated with the database 602 and to manage functions and sessions associated with the computing devices 612, 614. The server 604 can be associated with one or more cameras 608 of a system 620 for camera operation in a live environment 622.
[0068]
[0075] The encoder / decoder 618 may be associated with the system 620 to enable remote monitoring and control of aspects of camera operation in a live environment. The encoder / decoder 618 and the transmitter / receiver 616 may include a processor and a memory having instructions that, when executed by the processor, cause the encoder / decoder 618 and the transmitter / receiver 616 to collectively perform the encoding / decoding and transmitting / receiving functions described throughout this specification and with reference to at least Figures 2, 3, and 5. For example, camera control may include signals for camera selection and adjustment directly from the server 604 to the camera 608. However, an autostaging module may be enabled closer to the camera via the encoder / decoder 618 and the transmitter / receiver 616 to provide camera control of the camera 608.
[0069]
[0076] The server 604 can be in the live environment, or can be located separately from the live environment to perform the decoding functions described throughout this specification and with reference to at least Figures 2, 3, and 5. Such a server 604 can support a system 620 for camera operation in the live environment 622. Such a system 620 can operate partially within the live environment 622. Such a tool 620 can include subsystems for performing the functions described throughout this specification.
[0070]
[0077] The subsystems may be modules that may be capable of testing or training portions of the system 620, as described herein at least with respect to aspects of ODTW and tagging. The subsystems may be contained within one or more computing devices having at least one processor and memory, whereby the at least one processor performs functions based in part on instructions from the memory executed on the at least one processor. While shown collectively, the system boundary 620 may be part of one or more cameras 608 and encoders 618 and transmitters / receivers 616. The server 604 and computing devices 610-614 may be in various geographic locations other than the live environment.
[0071]
[0078] One or more cameras 608 of a system 620 for camera operation in a live environment are provided to enable capturing views associated with the live environment 622. The system 620 for camera operation in a live environment can be adapted to transmit the received information and further processed information via wired or wireless communication. The encoder and transmitter 616 can communicate with one or more components within the system 620 and external to the system 620.
[0072]
[0079] One or more of the components 604A-N can be adapted to function as provisioning devices within the server 604. Additionally, one or more of the components 604A-N can include one or more processors and one or more memory devices adapted to function for camera operation in a live environment, and other processors and memory devices in the server 604 can perform other functions.
[0073]
[0080] Server 604 may also provide software-based services or applications in a virtual or physical environment (such as to support the simulations referenced herein). When server 604 is a virtual environment, components 604A-N are software components that may be implemented in the cloud. This feature allows for remote access to information received and communicated between any of the aforementioned devices. One or more components 604A-N of server 604 may be implemented in hardware or firmware, in addition to the software implementations described throughout this specification. The methods and systems herein also allow for combinations thereof to be used.
[0074]
[0081] One of the computing devices 610-614 may be a smart monitor or display having at least a microcontroller and memory with instructions that enable display of information from the camera 608. One of the computing devices 610 may be a transmitter device for transmitting directly to a receiver device or for transmitting over the network 606 to a receiver device that may be part of an encoder and transmitter 616, and for transmitting to a server 604 or other computing devices 612, 614.
[0075]
[0082] The other computing devices 612, 614 may include portable handheld devices, including but not limited to smartphones, cellular phones, tablet computers, personal digital assistants (PDAs), and wearable devices (head-mounted displays, watches, etc.) Additionally, the other computing devices 612, 614 may run one or more operating systems, including Microsoft Windows Mobile®, Windows® (any generation), and / or various mobile operating systems such as iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®, and / or variants thereof.
[0076]
[0083] The other computing devices 612, 614 may support applications designed for Internet-related applications, electronic mail (email), short or multimedia message service (SMS or MMS) applications, and may use other communication protocols. The other computing devices 612, 614 may also include general-purpose personal computers and / or laptop computers running operating systems such as Microsoft Windows®, Apple Macintosh®, and / or Linux®. The other computing devices 612, 614 may also be workstations running UNIX® or UNIX-like operating systems, or other GNU / Linux operating systems such as Google Chrome OS®. Thin client systems, including gaming systems (such as Microsoft Xbox®), may also be used as the other computing devices 612, 614.
[0077]
[0084] The methods or systems herein employ network(s) 606, which can be any type of network capable of supporting data communications using various protocols, including TCP / IPT (transmission control protocol / Internet protocol), SNA (systems network architecture), IPX (Internet packet exchange), AppleTalk, and / or variations thereof. The network(s) 606 can be an Ethernet, token ring, wide area network, the Internet, a virtual network, a virtual private network (VPN), a local area network (LAN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (such as one operating using guidelines from organizations such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, Bluetooth, and / or any other wireless protocol), and / or a network based on any combination of these and / or other networks.
[0078]
[0085] Server 604 runs a suitable operating system, including any of the operating systems described throughout this specification. Additionally, server 604 also runs several server applications, including a hypertext transport protocol (HTTP) server, a file transfer protocol (FTP) server, a common gateway interface (CGI) server, a Java server, a database server, and / or variants thereof. Database 602 is supported by database server functionality in server 604 that provides front-end capabilities. Such database server functionality includes those from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variants thereof.
[0079]
[0086] The server 604 can provide real-time updates regarding the feed and / or media feed. The server 604 is part of a network of server boxes that are distributed over a range but function for the presently described process. The server 604 includes applications for measuring network performance through network monitoring and traffic management. A database 602 is provided that allows for the storage of information from the live environment, including user interactions, usage pattern information, adaptation rule information, and other information.
[0080]
[0087] The methods or systems herein are software as a service (SaaS) approaches that use instructions in a non-transitory medium that, when installed and executed on at least one processor, cause the processor of a system for camera operation in a live environment to perform a specific function. Creating video from content can be a function that ties together one or more aspects described herein, for example. The approaches herein can replace entire video production teams. At the same time, the approaches herein ensure higher standards in the industry for quality and artistic integrity.
[0081]
[0088] Organizers of classical music concerts and other types of concerts and live performances may use such methods and systems via web-based or standalone platforms. This gives event organizers complete control and flexibility for recording and / or streaming each live concert, at a minimum. Furthermore, the AI / ML approach herein can automatically generate a combination of live versions and multiple video streams according to predefined and changing scenarios (e.g., fixed renditions, changing renditions, adjustments, etc.). Thus, the video is a choreographed concert video. The system herein also ensures that every classical music performance is recorded and that every recording reflects the carefully planned rendition of the live environment.
[0082]
[0089] The methods and systems described herein provide concert organizers with access to tools that enable them to expand their activities in the digital and Internet realms. For example, the software approach described herein automates the process of recording a live environment (including the live version of the score in performance) by tracking the score against the recorded version. Predetermined and dynamic concert renditions can be achieved, such as by directing camera movements to performers or sections in accordance with the score and the conductor's vision. Furthermore, the methods and systems described herein ensure that the solutions are universally applicable to standard concerts, as well as to progressive (modern) symphony orchestras, music universities, conservatories, opera houses, and individual artists, where the high costs of recording can be reduced.
[0083]
[0090] The methods and systems described herein can use ODTW to match a live version of a musical score to a recorded version. In a live environment, ODTW can determine in real time the points in the live version of a musical score relative to the recorded version. Thus, using a previously prepared script, it is possible to control the direction of one or more cameras for recording or displaying video related to the dominant influence of the live version. For this purpose, a camera is used that allows remote control of the lens and saving of the frames used. Thus, the camera simply broadcasts or transmits according to the dominant influence, which may be tagged as an instruction in the recorded version. Furthermore, images from the cameras can be recorded in multi-track and mixed formats, showing images from one camera according to the tag.
[0084]
[0091] The methods and systems herein incorporate software with a set of libraries that enable, among other things, many of the features described herein. For example, one feature is tracking musical scores against tonal music standards and against standards in contemporary and atonal music. Furthermore, recording a live version, including a concert, according to a direction set by the concert organizer avoids the need for manual camera calibration and translation of high-level tags to specific camera settings. Furthermore, a mixture of multi-track camera recordings and synchronized live versions with video can be used, with imagery from one camera selected by an administrator. Thus, the live environment can be, at a minimum, a playback of a previously recorded live version, with the video portion altered based on additional camera footage that is available and selected or adjusted to fit one or more frames of the live version playback.
[0085]
[0092] The methods or systems herein allow for fully automatic rendition without the need for additional information by storing and validating tags in the recorded version. For example, such information related to dominant influences can be converted into high-level tags and provided for specific camera settings. Thus, using knowledge of the tags herein representing an established scene layout and techniques for automatic scene calibration (using tools such as depth cameras), it is possible to record a rendition for a multi-track camera recording, with audio and video mixed and synchronized, always showing images from one selected camera.
[0086]
[0093] The methods or systems herein may be based on an audio recognition problem, where the question of where a performer is currently in a musical score can be solved. Furthermore, the sound in the current live version can be adjusted relative to the recorded version. The recorded version can be assigned or determined by an administrator, such as by direction, using a stream plan. The determination must be annotated in the system for camera operation in a live environment to enable the methods herein, and thus the determination of the musical score to be performed is an action within the system, such as a computer-based input that initiates the functions described herein for camera operation in a live environment.
[0087]
[0094] The method or system described herein can match a live version to a recorded version using ODTW, which can operate on appropriately processed sound (e.g., using chroma features, variational autoencoders, vector quantization, and variational autoencoders). Furthermore, in a live environment, the ODTW module can determine in real time which part of the score is being played, and such determination can be used to control camera movement. The camera allows for remote lens movement and frame saving unless currently in record mode. Furthermore, images from the cameras can be recorded in a multi-track and sound-video synchronized fashion, always showing images from one selected camera according to a prepared scenario.
[0088]
[0095] The systems and methods described herein include software that runs on a server located remotely rather than in the live environment. For example, parts of the software can be deployed in the cloud, allowing a client panel to make various adjustments. Furthermore, retransmission of audio / video signals can occur after they are received from an on-site server in the live environment to a remote environment. The systems and methods described herein are scalable to accommodate different numbers of cameras, allowing the system for camera operation in live environments to be adapted to various customer requirements.
[0089]
[0096] The method or system herein allows for the tracking of classical music tracks and the recording and / or broadcasting of live environments such as live concerts. There may be necessary programming layers and data flow between the live and recorded versions of the score. There may be a key layer supported by an innovative programming library responsible for pre-processing the live version for tonal music tracking or novel processing methods using variational autoencoder or VQ-VAE neural networks for contemporary and atonal music tracking.
[0090]
[0097] The methods and systems herein may use online dynamic time warping (ODTW) to determine the position of a live version relative to a recorded version during a live performance. The methods and systems herein allow tracking systems to be overhauled based on multiple recorded versions of the same musical score. This element stabilizes the operation of the algorithms described herein with respect to ODTW and AI / ML aspects, ensuring the correct process of tracking tracks that may have different interpretations and tempos.
[0091]
[0098] The automatic scene calibration feature can be fully automatic rendition without additional information in the form of translating high-level tags to specific camera settings. Additionally, in the methods and systems herein, there can be remaining components that accommodate programming requests typically found in an organization's infrastructure. Such requests can include scheduling recorded versions and tracking live versions according to a schedule.
[0092]
[0099] The systems and methods herein rely on programming libraries that implement methods for processing audio signals, that support determining the position of a live version relative to the position of a recorded version and that implement an On-Demand Two (ODTW) approach, that support a process for updating the systems herein for camera operation in a live environment, that implement tools for stage preparation with particular emphasis on determining camera movement relative to the recorded version, and that support a process for mapping labels to camera movement with particular emphasis on remote movement of the camera lens. Overall, the frames and their progression can be saved according to a sequence that is the output of the libraries that implement tools for camera operation in a live environment.
[0093]
[0100] Additionally, the systems and methods herein include a programming library for implementing methods for recording either multi-track or mixed-track recordings, an advancement library for implementing tools for synchronizing a live version with a mixed recorded version, and a programming library for implementing methods for automatic scene calibration.
[0094]
[0101] Furthermore, the programming libraries herein can be adapted to determine the type of musicians (instruments, etc.), to implement methods for determining the positions (coordinates) of people on stage (using neural networks for detection) based in part on the calibrated space specified by the library for implementing tools for automatic scene calibration, and then comparing those coordinates with the coordinates of the musicians' sections in an inaccurate stage plan.
[0095]
[0102] Further programming libraries here support directing the preparation process. There is a dedicated general pattern or standard for the software architecture used in the method and system. This approach takes into account the necessary programming layers and the data / information flowing between them. Additionally, there may be key layers corresponding to libraries / modules.
[0096]
[0103] The methods or systems herein include instructions (including documentation and sample realization) for integrating programming layers within the above-referenced features, taking into account present and missing components. The systems and methods herein enable the recording and / or transmission of live environments (including live performances) using multiple cameras mixed in real time, addressing the complexity and cost of coordinating the recording and / or broadcast of live concerts handled by professional live production and production companies.
[0097]
[0104] Additionally, many people may be involved in a live production, including camera operators, art directors, and score readers, among others. There may be costs associated with managing streaming in terms of technicians and equipment. There may be logistical concerns with such live productions, such as those associated with recording from multiple cameras.
[0098]
[0105] The systems and methods herein also enable concert organizers and individual artists to digitally participate and share their concerts online, making it easier for them to reach a wider audience. Implementing such a goal with the developed solution would require using professional companies to handle live broadcast production and implementation, resulting in complexity and cost that exceeds the benefits that would be gained by those seeking such content. Automating the entire process of broadcasting a live classical music concert using the developed system would allow concert organizers to reduce costs and complexity while simultaneously increasing the potential audience.
[0099]
[0106] The systems and methods herein can serve as a unified camera operator, score reader, and art director. The systems and methods address issues associated with atonal classical music using preprocessing with neural network models supported by the VQ-VAE approach described with respect to one or more figures herein. The systems and methods herein can make high-quality multi-camera concert broadcasts accessible to everyone, making classical music more accessible than ever before. Furthermore, such an approach can replace previous concert models that required extensive planning and high costs. A non-limiting example of the systems and methods can enable customers to quickly and easily record each performance.
[0100]
[0107] A method is provided herein for tracking a live version of a musical score to enable camera operation in a live environment. A system associated with the method includes a memory and at least one processor for executing instructions stored in the memory to perform the steps of the method. The steps of the method include determining a musical score to be performed in the live environment. The method further includes analyzing the live version of the musical score with a recorded version of the musical score in the live environment. The analysis is for determining a time-based difference of at least one frame of the live version relative to the recorded version, and the time-based difference, along with at least one dominant influence of the recorded or live version, is used to enable camera selection or camera adjustment of one or more cameras in the live environment.
[0101]
[0108] Another method herein is for camera operation in a live environment. A system associated with the method includes a memory and at least one processor for executing instructions stored in the memory to perform the steps of the method. The steps of the method include determining a musical score to be performed in the live environment. The method further includes determining at least one dominant influence in a live or recorded version of the musical score. The live version of the musical score is analyzed frame-by-frame with the recorded version of the musical score in the live environment. The method includes using time-based differences from the frame-by-frame analysis to select or adjust one or more cameras in the live environment to point in a direction associated with the at least one dominant influence.
[0102]
[0109] The technology herein is susceptible to modifications and alternative constructions, which variations are within the spirit of the disclosure. Accordingly, while certain example embodiments are shown in the drawings and described above in detail, it is not intended to limit the disclosure to the particular form or forms disclosed, but rather to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure as defined by the appended claims.
[0103]
[0110] In the context of describing the disclosed embodiments (particularly in the context of the claims below), terms such as a, an, the, and similar references are understood to include both the singular and the plural and are not intended to define the terms, unless otherwise indicated herein or clearly contradicted by context. Including, having, including, and containing are understood to be open-ended terms (meaning phrases such as, including, but not limited to, etc.) unless otherwise noted. Connected, when unmodified and referring to a physical connection, can be understood as being partially or wholly contained within, attached to, or joined together, even if there is something intervening.
[0104]
[0111] The recitation of ranges of values herein, unless otherwise indicated herein, is intended to be used merely as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated into the specification as if it were individually recited herein. The use of terms such as set (with respect to a set of items) or subset, etc., is to be understood as a non-empty collection containing one or more members, unless otherwise indicated herein or contradicted by context. Furthermore, unless otherwise indicated herein or contradicted by context, the term subset of a corresponding set does not necessarily indicate a proper subset of the corresponding set, and a subset and a corresponding set may be equivalent.
[0105]
[0112] Conjunctive language, such as phrases of the form "at least one of A, B, and C," or "at least one of A, B, and C," is understood in its commonly used context to indicate that an item, term, etc. can be either A or B or C, or any non-empty subset of the set A, B, and C, unless otherwise indicated herein or clearly contradicted by context. Furthermore, conjunctive language for a set having three members, such as "at least one of A, B, and C," or "at least one of A, B, and C," refers to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}. Thus, such conjunctive language is not generally intended to imply that a particular embodiment requires that at least one of A, at least one of B, and at least one of C are each present. Further, unless otherwise stated or contradicted by context, terms such as "plurality," "a plurality," and the like refer to a plurality (e.g., "a plurality of items" refers to a number of items). A number of items in a plurality means at least two, but may be so if explicitly or the context indicates a greater number. Further, unless stated otherwise or clear from context, phrases such as "based on," "based on," and the like mean "based at least in part on," not "based only on."
[0106]
[0113] 5A and 5B and the sub-steps described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. A method or system herein includes processes such as those described herein (or variations and / or combinations thereof), which may be performed under the control of one or more computer systems configured with executable instructions, and which may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed collectively or exclusively by one or more processors in hardware or a combination thereof.
[0107]
[0114] Such code can be stored on a computer-readable storage medium. Such code can be a computer program having instructions executable by one or more processors. A computer-readable storage medium is a non-transitory computer-readable storage medium, which excludes transitory signals (such as propagating momentary electrical or electromagnetic transmissions), but includes non-transitory data storage circuitry (such as buffers, caches, and queues) within a transmitter or receiver of a transitory signal. Furthermore, code (executable code or source code) can be stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) that contain executable instructions, which, when executed by (e.g., as a result of) one or more processors of a computer system, cause the computer system to perform the operations described herein.
[0108]
[0115] A set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, where one or more of the individual non-transitory computer-readable storage media lack all of the code, and the multiple non-transitory computer-readable storage media collectively store all of the code. The executable instructions are executed such that different instructions are executed by different processors, where the non-transitory computer-readable storage media store the instructions, a main central processing unit (CPU) executes some of the instructions, and other processing units execute other instructions. Furthermore, various components of the computer system have separate processors, where the various processors execute different subsets of the instructions.
[0109]
[0116] A computer system may be configured to implement one or more services that, individually or collectively, perform the operations of the processes described herein, and such a computer system may be configured with applicable hardware and / or software that enables the operations to be performed. A computer system that implements at least one embodiment of the present disclosure may be a single device or a distributed computer system having multiple devices that perform different operations, where a distributed computer system performs the operations described herein but where no single device performs all of the operations.
[0110]
[0117] While the above discussion provides at least one embodiment for implementing the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. Additionally, while certain responsibilities may be distributed among components and processes, which are defined above for purposes of discussion, various functions and responsibilities may be distributed and divided in various ways depending on the environment.
[0111]
[0118] Although the subject matter has been described in specific language regarding structures and / or methods or processes, it is understood that the subject matter claimed in the appended claims is not limited to the particular structures or methods described. Instead, the particular structures or methods are disclosed as examples of how the claims could be implemented.
[0112]
[0119] From all of the above, those skilled in the art will readily appreciate that the tools of the present disclosure provide many technical and commercial advantages and can be used in a variety of applications. It will be readily understood that various embodiments can be combined or modified based in part on the present disclosure, and that the present disclosure supports such combinations and modifications to achieve the benefits described above.
Claims
1. 1. A processor-implemented method, comprising: determining, by at least one processor, that at least one frame of a live version of a musical score being performed in a live environment corresponds to a recorded frame of a recorded version of said musical score; enabling, by at least one camera, a camera movement according to a time-based reference between said at least one frame and said recorded frame and according to a dominant influence of at least one of said recorded version or said live version; A method comprising:
2. 10. The processor-implemented method of claim 1, comprising: the dominant influence is in the form of instructions stored in memory along with the recorded version of the musical score and executed by the at least one processor, the instructions including causing at least one of the musical performers to be captured by the at least one camera using the camera movement for at least one frame, or to be captured by the at least one camera in such a way that a type of immersive experience is created, the processor-implemented method comprising: enabling camera movement to capture the musical performer or the immersive experience; causing the camera movement to be included in the live version of the live environment; The method further comprises:
3. 3. The processor-implemented method of claim 2, comprising: The method further comprising the step of enabling the instructions to be provided in a computer-based input and based on the recorded version of the musical score.
4. 3. The processor-implemented method of claim 2, comprising: The method, wherein the immersive experience is generated for the at least one frame of the live version, and the type of immersive experience is at least one of a graphical visualization being displayed during a live performance and text being displayed to an audience in the live environment.
5. 5. The processor-implemented method of claim 4, comprising: The method further includes adjusting the camera movement or the immersive experience in response to features derived from audio and video of the live environment, the features including at least one of a volume and a level of emotional expression of the musical performer, and the adjusting step including at least one of zooming a camera and adjusting a color of the graphical visualization or the text.
6. 6. The processor-implemented method of claim 5, comprising: The method further comprising determining the level of emotional expression of the musical performer based at least in part on a machine learning algorithm that tracks a facial matrix of the musical performer in the live version of the musical score.
7. 6. The processor-implemented method of claim 5, comprising: A method wherein adjusting the camera movement or the immersive experience is further based in part on object tracking performed with respect to the live version of the musical score.
8. 10. The processor-implemented method of claim 1, comprising: The method, further comprising the step of enabling camera movement with a plurality of cameras of the at least one camera, the camera movement including a particular action performed on a particular one of the plurality of cameras.
9. 10. The processor-implemented method of claim 8, comprising: The method further comprising modifying the camera movement in real time based at least in part on determining a problem with the at least one camera or a further camera associated with the at least one camera.
10. 10. The processor-implemented method of claim 1, comprising:
10. The method of claim 9, further comprising: determining, based on a machine learning algorithm, the recorded frames of the recorded version of the musical score relative to the at least one frame of the live version of the musical score, wherein the machine learning algorithm extracts features describing the at least one frame and the recorded frames and uses online dynamic time warping (ODTW) to determine associations between the features of the audio.
11. 11. The processor-implemented method of claim 10, continuously improving the machine learning algorithm based at least in part on feedback provided by a user, the feedback being collected via computer-based input, the feedback including errors in matching the frames; the feedback is added to a training data set for the machine learning algorithm. method.
12. 11. The processor-implemented method of claim 10, The method further comprises the step of enabling the machine learning algorithm to use results from a previous live environment of the musical score, wherein the recorded frame determined using the at least one frame enables consecutive frame matches between the live and recorded versions.
13. 11. The processor-implemented method of claim 10, The method further comprises the step of enabling the machine learning algorithm to use a confidence level, wherein a first predetermined value of the confidence level allows the machine learning algorithm to influence the camera movement by one or more of: switching to a dominant influence of another type or setting; and maintaining the at least one camera or multiple cameras in a switched mode until the machine learning algorithm regains a second predetermined value of the confidence level for a next frame of the live or recorded version of the musical score.
14. 14. The processor-implemented method of claim 13, comprising: The method, wherein the switching includes at least one of directing the at least one camera or cameras to capture a conductor of the musical score based on an immersive experience associated with the live version of the musical performance, and reducing an influence associated with the at least one dominant influence.
15. 10. The processor-implemented method of claim 1, comprising: storing a plurality of recorded versions of the musical score together with associated dominant influences; determining a recorded version from a plurality of recorded versions to be used with the live version at the start of the score based in part on differences between the at least one frame and respective ones of a plurality of recorded frames of the plurality of recorded versions; The method further comprises:
16. 10. The processor-implemented method of claim 1, comprising:
10. The method of claim 9, further comprising determining a recorded version of a plurality of recorded versions to be used with the live version of the musical score based in part on a confidence level, wherein a machine learning algorithm responsible for determining the recorded frame associated with the at least one frame according to the confidence level provides an output using the at least one frame for the plurality of recorded frames.
17. 10. The processor-implemented method of claim 1, comprising: enabling the camera movement for the at least one camera or cameras using preliminary instructions until the recorded version is determined using a predetermined value of confidence, wherein the determination of the recorded version is performed in real time or near real time using a scalable cloud environment and using multiple recorded versions; sending information associated with said recorded version to said at least one camera or cameras; using an edge environment associated with the scalable cloud environment to enable further real-time determination of additional recorded frames to be used with the live version; The method further comprises:
18. 10. The processor-implemented method of claim 1, comprising: The method further includes capturing the live version with the at least one camera or multiple cameras, the capturing being synchronized with live audio in real time, for use in at least one of online broadcasting, television-based live entertainment production, and sheet music recording production.
19. 20. The processor-implemented method of claim 18, comprising: The method further includes the step of enabling camera movements for various of the online broadcast, the television-based live entertainment production, and the production of the score recording to obtain multiple live versions of the score.
20. 10. The processor-implemented method of claim 1, further comprising the step of enabling at least one function associated with the processor-implemented method to be turned on or off using a computer-based input reporting at least one aspect of the musical score, the at least one function being to control a level of automation of camera movement.
Citation Information
Patent Citations
Imaging system and control method thereof, and computer program
JP2017028585A
Control device, control method and program
JP2019140642A
Network-based processing and delivery of multimedia content for live music performances
JP2019525571A
Processor and program
JP2021060873A
Video data processing device, video distribution system, video editing device, recording medium, video data processing method, video distribution method, and program
JP2021153257A