Methods for determining camera movement information, methods and devices for displaying scene images

By analyzing the background music file to determine the timing of camera movement transitions and switching the camera movement information of the virtual camera, the problem of determining the camera movement information of the virtual camera is solved, thus improving the performance of the virtual scene.

CN116489451BActive Publication Date: 2025-10-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310158257.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-10-31
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

In virtual technology, how to determine the camera movement information of a virtual camera to improve the performance of scenes shot in virtual scenes, especially when switching camera movement information at the turning points of background music.

Method used

By analyzing the background music file, multiple target turning points on the playback timeline are identified, and the camera movement switching times are determined based on these turning points. Target camera movement information is then selected, allowing the virtual camera to switch camera movement information at these times.

Benefits of technology

It enables switching the camera movement information of the virtual camera at the turning points of the background music, thereby improving the performance of the scene images obtained by the virtual camera shooting virtual scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489451B_ABST
    Figure CN116489451B_ABST
Patent Text Reader

Abstract

This application discloses a method for determining camera movement information, a method for displaying scene images, and an apparatus, belonging to the field of computer technology. The method includes: acquiring a background music file and multiple candidate camera movement information, where the background music is music accompanying the display of scene images, and the scene images are images obtained by capturing a virtual scene using a virtual camera; parsing the background music file to obtain multiple target turning points on the background music playback timeline; determining multiple camera movement switching moments on the playback timeline based on the multiple target turning points; for any given camera movement switching moment, selecting target camera movement information from the multiple candidate camera movement information, whereby the target camera movement information is used to enable the virtual camera to start capturing the virtual scene from any given camera movement switching moment. This application enables switching the virtual camera's camera movement information at the turning points of the background music, improving the performance of the scene images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for determining camera movement information, a method for displaying scene images, and an apparatus. Background Technology

[0002] With the development of computer technology, virtual technology has also developed and is widely used in industries such as live streaming, animation, and gaming. Virtual technology involves constructing virtual scenes, using virtual cameras to capture images of these scenes, and then displaying these images on electronic devices.

[0003] During the filming of virtual scenes, the position and parameters of the virtual camera are constantly changing. This information, including the position and parameters of the virtual camera, is also known as the camera movement information. Therefore, determining the camera movement information of the virtual camera becomes a crucial problem that needs to be solved. Summary of the Invention

[0004] This application provides a method for determining camera movement information, a method for displaying scene images, and a device that can switch the camera movement information of a virtual camera at the turning points of background music, thereby improving the performance of scene images obtained by the virtual camera shooting virtual scenes. The technical solution includes the following contents.

[0005] Firstly, a method for determining camera movement information is provided, the method comprising:

[0006] Obtain the background music file, which is the music that accompanies the display of scene images, which are images obtained by capturing virtual scenes with a virtual camera;

[0007] Acquire multiple candidate camera movement information, where any candidate camera movement information is information from the virtual camera.

[0008] The background music file is parsed to obtain multiple target turning points on the playback timeline of the background music;

[0009] Based on the multiple target turning points, multiple camera switching moments on the playback timeline are determined;

[0010] For any camera movement switch moment, a target camera movement information is selected from the plurality of candidate camera movement information. The target camera movement information is used to enable the virtual camera to start shooting the virtual scene from the any camera movement switch moment based on the target camera movement information.

[0011] Secondly, a method for displaying a scene is provided, the method comprising:

[0012] The live stream interface is displayed and background music is played. The live stream interface includes a first scene screen, which is a screen obtained by shooting a virtual scene with a virtual camera based on first camera movement information.

[0013] In response to the background music playing at the moment of camera movement switching, the first camera movement information is switched to the second camera movement information;

[0014] The live streaming interface displays a second scene, which is a shot of the virtual scene taken by the virtual camera based on the second camera movement information.

[0015] Thirdly, a device for determining camera movement information is provided, the device comprising:

[0016] The acquisition module is used to acquire background music files. The background music is the music that accompanies the display of scene images, which are images obtained by capturing virtual scenes with a virtual camera.

[0017] The acquisition module is also used to acquire multiple candidate camera movement information, where any one of the candidate camera movement information is information of the virtual camera.

[0018] The parsing module is used to parse the background music file and obtain multiple target turning points on the playback timeline of the background music.

[0019] The determination module is used to determine multiple camera movement switching moments on the playback timeline based on the multiple target turning points;

[0020] The selection module is used to select target camera information from the plurality of candidate camera information at any camera movement switching moment. The target camera information is used to enable the virtual camera to start shooting the virtual scene from the any camera movement switching moment based on the target camera information.

[0021] In one possible implementation, the acquisition module is configured to, for any candidate camera movement information, acquire position information of multiple trajectory points on the motion trajectory of the virtual camera; determine the motion trajectory of the virtual camera based on the position information of the multiple trajectory points; set motion information of a reference object on the motion trajectory to obtain the any candidate camera movement information, wherein the reference object is used to make the virtual camera follow the reference object and move on the motion trajectory according to the motion information of the reference object.

[0022] In one possible implementation, the parsing module is used to extract audio information from the background music file; perform content detection processing on the audio information to obtain multiple audio turning points on the playback timeline; and determine multiple target turning points based on the multiple audio turning points.

[0023] In one possible implementation, the parsing module is used to segment the audio information to obtain multiple first audio segments; extract features from the multiple first audio segments to obtain features of each first audio segment; classify the multiple first audio segments based on the features of each first audio segment to obtain the beat type to which each first audio segment belongs; and determine the multiple audio inflection points based on the beat type to which each first audio segment belongs.

[0024] In one possible implementation, the parsing module is used to segment the audio information to obtain multiple second audio segments; extract features from the multiple second audio segments to obtain features of each second audio segment; perform clustering processing on the features of each second audio segment to obtain at least one cluster, where each cluster includes at least one second audio segment; and determine the multiple audio inflection points based on the at least one cluster.

[0025] In one possible implementation, the parsing module is used to extract lyrics from the background music file; perform sentiment analysis on the lyrics to obtain multiple emotional turning points on the playback timeline; and determine multiple target turning points based on the multiple emotional turning points.

[0026] In one possible implementation, the parsing module is used to classify the lyrics text to obtain global sentiment information of the lyrics text and local sentiment information of each lyric phrase included in the lyrics text; for any lyric phrase, based on the global sentiment information and the local sentiment information of the lyric phrase, an sentiment-related index corresponding to the lyric phrase is determined, and the sentiment-related index corresponding to the lyric phrase is used to characterize the degree of correlation between the local sentiment information of the lyric phrase and the global sentiment information; based on the sentiment-related index corresponding to each lyric phrase, the plurality of sentiment turning points are determined.

[0027] In one possible implementation, the selection module is used to randomly generate a target camera movement number, select the target camera movement information corresponding to the target camera movement number from the plurality of candidate camera movement information, wherein each candidate camera movement information corresponds to a camera movement number; or, based on the operation of selecting any candidate camera movement information in the camera movement display list, the any candidate camera movement information is determined as the target camera movement information, and the camera movement display list is used to display the plurality of candidate camera movement information.

[0028] Fourthly, a display device for a scene image is provided, characterized in that the device comprises:

[0029] The display module is used to display the live broadcast interface, which includes a first scene image. The first scene image is a picture obtained by a virtual camera capturing a virtual scene based on first camera movement information.

[0030] The playback module is used to play background music;

[0031] The switching module is used to switch the first camera movement information to the second camera movement information in response to the background music playing at the camera movement switching moment.

[0032] The display module is also used to display a second scene on the live streaming interface. The second scene is a picture of the virtual scene captured by the virtual camera based on the second camera movement information.

[0033] In one possible implementation, the switching module is configured to display a camera movement switching control on the live streaming interface; and, based on a confirmation operation on the camera movement switching control, switch the first camera movement information to the second camera movement information.

[0034] Fifthly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by the processor to enable the electronic device to implement the method for determining camera movement information as described in the first aspect or the method for displaying scene images as described in the second aspect.

[0035] In a sixth aspect, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to enable an electronic device to implement the method for determining camera movement information as described in the first aspect or the method for displaying scene images as described in the second aspect.

[0036] In a seventh aspect, a computer program or computer program product is also provided, wherein the computer program or computer program product stores at least one computer program, which is loaded and executed by a processor to enable an electronic device to implement the method for determining camera movement information as described in the first aspect or the method for displaying scene images as described in the second aspect.

[0037] The technical solution provided in this application brings at least the following beneficial effects:

[0038] The technical solution provided in this application involves parsing the background music file to obtain multiple target turning points on the playback timeline, determining multiple camera movement switching moments on the playback timeline based on each target turning point, and identifying the target camera movement information corresponding to each camera movement switching moment. By determining the target camera movement information corresponding to the camera movement switching moments, the camera movement information of the virtual camera can be switched at the turning points of the background music, thereby improving the performance of the scene images obtained by the virtual camera shooting the virtual scene. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a schematic diagram of the implementation environment of a method for determining camera movement information provided in an embodiment of this application;

[0041] Figure 2 This is a flowchart of a method for determining camera movement information provided in an embodiment of this application;

[0042] Figure 3 This is a schematic diagram of a display page provided in an embodiment of this application;

[0043] Figure 4 This is a flowchart illustrating a method for displaying a scene image according to an embodiment of this application;

[0044] Figure 5 This is a schematic diagram illustrating the determination of camera movement information provided in an embodiment of this application;

[0045] Figure 6 This is a schematic diagram illustrating the determination of camera movement switching time according to an embodiment of this application;

[0046] Figure 7 This is a schematic diagram of the structure of a camera movement information determination device provided in an embodiment of this application;

[0047] Figure 8 This is a schematic diagram of the structure of a scene display device provided in an embodiment of this application;

[0048] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0049] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0051] Figure 1 This is a schematic diagram illustrating the implementation environment of a method for determining camera movement information provided in an embodiment of this application. Figure 1 As shown, the implementation environment includes a terminal device 101 and a server 102. The method for determining camera movement information in this embodiment can be executed by the terminal device 101, the server 102, or jointly by both.

[0052] Terminal device 101 can be a smartphone, game console, desktop computer, tablet computer, laptop computer, smart TV, smart in-vehicle device, smart voice interaction device, smart home appliance, etc. Server 102 can be a single server, a server cluster consisting of multiple servers, or any of the following: cloud computing platform and virtualization center. This application embodiment does not limit this. Server 102 can communicate with terminal device 101 via a wired network or wireless network. Server 102 can have functions such as data processing, data storage, and data transmission and reception. This application embodiment does not limit this. The number of terminal devices 101 and servers 102 is not limited and can be one or more.

[0053] In industries such as live streaming, animation, and gaming, virtual scenes can be constructed, virtual cameras can be used to capture images of these virtual scenes, and then electronic devices can display these images.

[0054] During the filming of virtual scenes, the position and parameters of the virtual camera are constantly changing. This information, including the virtual camera's position and parameters, is also known as the virtual camera's camera movement information. Determining this camera movement information is a crucial problem that needs to be solved.

[0055] This application provides a method for determining camera movement information. This method can be applied to the aforementioned implementation environment and can switch the camera movement information of the virtual camera at the turning points of the background music, thereby improving the performance of the scene images obtained by the virtual camera shooting virtual scenes. Figure 2 The flowchart shown in this application embodiment illustrates a method for determining camera movement information. For ease of description, the terminal device 101 or server 102 executing the method for determining camera movement information in this application embodiment is referred to as an electronic device. This method can be executed by an electronic device. Figure 2 As shown, the method includes the following steps.

[0056] Step 201: Obtain the background music file and multiple candidate camera movement information.

[0057] In virtual technology, 3D modeling techniques can be used to construct virtual scenes so that electronic devices can access them. This application's embodiments do not elaborate on the construction method of the virtual scene. The virtual scene includes at least one of static object models and dynamic object models. Static object models are stationary object models; for example, building models and artificial mountain models are static object models. Dynamic object models are object models that can change their posture, position, etc.; for example, biological models are dynamic object models.

[0058] Typically, a virtual camera needs to be configured. A virtual camera is a simulation of a real-world camera in the computer world. The parameters of a virtual camera include intrinsic and extrinsic parameters. Intrinsic parameters include at least one of the following: camera focus, focal length, target area, distortion parameters, etc. Extrinsic parameters include at least one of the following: camera rotation information, displacement information, etc.

[0059] A virtual camera is used to capture images of a virtual scene, which are then displayed on an electronic device. Typically, background music plays during this display, and the device can access the music file. In other words, the scene images are captured by a virtual camera, and the background music accompanies the display.

[0060] This application does not limit the way the electronic device obtains background music files. For example, the electronic device may use any music file captured from the network as the background music file, or the electronic device may use the music file input by the object as the background music file, or the electronic device may use any music file read from the storage device and stored in the storage device as the background music file.

[0061] The electronic device can also acquire multiple candidate camera movement information. Each candidate camera movement information is information about a virtual camera, including the virtual camera's parameters and motion trajectory.

[0062] In one possible implementation, step 201, "acquiring multiple candidate camera movement information," includes steps 2011 to 2013 (not shown in the figure).

[0063] Step 2011: For any candidate camera movement information, obtain the position information of multiple trajectory points on the motion trajectory of the virtual camera.

[0064] The electronic device is equipped with multiple trajectory generation functions. Any one of these functions is a line function, which can be any of the following: a straight line function, a polyline function, an ellipse function, a curve function, etc. The electronic device can randomly select any one of these functions and use it as the target function. Alternatively, the electronic device can select a trajectory generation function from among the multiple functions based on an object, and use the selected function as the target function. The target function is then used to calculate the position information of multiple trajectory points, where each trajectory point is a point on the motion trajectory of the virtual camera.

[0065] The position information of any trajectory point is used to represent its location in the virtual scene. The virtual scene corresponds to a three-dimensional coordinate system, and the three-dimensional coordinates of any trajectory point in the three-dimensional coordinate system can be used as the position information of that trajectory point.

[0066] Step 2012: Determine the motion trajectory of the virtual camera based on the position information of multiple trajectory points.

[0067] The position information of multiple trajectory points can be sorted to obtain the sorted position information of each trajectory point, and then stored in an index file. The trajectory points are then connected sequentially according to their sorted position information to obtain the motion trajectory of the virtual camera.

[0068] In one possible implementation, the sorted trajectory points are designated as first trajectory points. For any two adjacent first trajectory points, their position information can be interpolated to obtain the position information of at least one second trajectory point, where any second trajectory point lies between the two first trajectory points. In this way, the position information of the second trajectory points located between every two first trajectory points can be determined. Based on the position information of each first trajectory point and each second trajectory point, the first and second trajectory points are sequentially connected to obtain the motion trajectory of the virtual camera.

[0069] Step 2013: Set the motion information of the reference object on the motion trajectory to obtain any candidate camera movement information. The reference object is used to make the virtual camera follow the reference object and move on the motion trajectory according to the motion information of the reference object.

[0070] A reference object can be set on the motion trajectory. This reference object can be any object, such as a sphere or a plane. Next, motion information for the reference object is set on the motion trajectory. This motion information describes the reference object's motion along the trajectory and includes at least one of the following: velocity, acceleration, and direction. By setting the motion information of the reference object on the motion trajectory, candidate camera movement information is obtained. Since the virtual camera can follow the reference object and move along the trajectory according to its motion information, the virtual camera's motion on the trajectory is actually the motion of the reference object along the trajectory.

[0071] By selecting different trajectory generation functions and / or setting different motion information of reference objects on the motion trajectory, multiple candidate camera movement information can be determined in accordance with steps 2011 to 2013.

[0072] It is understandable that any candidate camera movement information may include not only the motion information of the reference object on the motion trajectory, but also the parameters of the virtual camera. The parameters of the virtual camera can be set before step 2011 to obtain any candidate camera movement information, or the parameters of the virtual camera can be set after "setting the motion information of the reference object on the motion trajectory" in step 2013 to obtain any candidate camera movement information.

[0073] When setting the parameters of the virtual camera, the camera's shooting object (i.e., the subject object, which in this embodiment is a virtual scene), camera intrinsic parameters, and camera extrinsic parameters can be configured through the virtual camera's configuration interface. For example, when configuring the camera's intrinsic parameters, the focus control object, camera focal length, and camera target surface information can be configured. The focus control object is the object focused on by the virtual camera's focus, generally the subject object or a part of the subject object.

[0074] Step 202: Parse the background music file to obtain multiple target turning points on the background music playback timeline.

[0075] The playback duration of the background music can be extracted from the background music file, and a playback timeline for the background music can be generated based on the playback duration. In an exemplary embodiment, the background music file includes a file header, frame files, and a file footer.

[0076] The file header, also called TAG_V2 (ID3V2), can be empty or include at least one of the following: the artist performing the background music, the name of the background music, the album to which the background music belongs, the release year of the background music, the genre of the background music, the duration of the background music, the sampling rate, the bit rate, and the encoding format.

[0077] The frame file includes at least one of an audio file and a lyrics (LRC) file. The audio file includes audio information, and the lyrics file includes lyrics text.

[0078] The file header, also called TAG_V1 (ID3V1), can be empty or include at least one of the following: the artist of the background music, the name of the background music, the album to which the background music belongs, the release year of the background music, the genre of the background music, the duration of the background music, the sampling rate, the bit rate, and the encoding format. Generally, the information in the file header is greater than or equal to the information in the file header.

[0079] The electronic device can directly read the playback duration of the background music from the file header or footer. Alternatively, it can determine the playback duration by parsing the frame file. Then, it generates a playback timeline based on the playback duration, with the corresponding duration being the playback length. Furthermore, based on the playback timeline, the electronic device parses the frame file to obtain multiple target inflection points located on the playback timeline. Each target inflection point is a turning point in target information on the playback timeline, including at least one of the following: the beat of the background music, the musical structure of the background music, and the emotional information expressed by the background music.

[0080] Since a frame file includes at least one of an audio file and a lyrics file, it can be parsed and processed according to any of the following implementations A through C.

[0081] In implementation A, the frame file includes an audio file. In this case, step 202 includes steps A1 to A3 (not shown in the figure).

[0082] Step A1: Extract audio information from the background music file.

[0083] Since the background music file includes frame files, and the frame files include audio files, the audio files can be extracted from the background music file, and the audio information contained in the audio files can be obtained.

[0084] Step A2 involves performing content detection processing on the audio information to obtain multiple audio turning points on the playback timeline.

[0085] The electronic device is equipped with a toolkit for content detection processing of audio information. By calling the toolkit to perform content detection processing on the audio information, multiple playback moments can be obtained, each of which is an audio inflection point on the playback timeline. This application does not limit the toolkit; for example, the toolkit may be Librosa or the Music Structure Analysis Framework (MSAF). Librosa is a toolkit for analyzing and processing audio and music, and can be used for beat detection processing of audio information. MSAF can be used for structure detection processing of audio information.

[0086] In one possible implementation, the content detection process is a beat detection process. In this case, step A2 includes steps A21 to A24 (not shown in the figure).

[0087] Step A21: The audio information is segmented to obtain multiple first audio segments.

[0088] The window size and sliding distance can be set. The audio information is segmented based on the window size and sliding distance to obtain multiple first audio segments. The playback duration of any first audio segment is the window size. On the playback timeline, the duration between the start playback times of two adjacent first audio segments is the sliding distance, and the duration between the end playback times of these two first audio segments is also the sliding distance.

[0089] Optionally, the window size is greater than or equal to the sliding distance. If the window size is greater than the sliding distance, then two adjacent first audio segments overlap. For example, if the window size is 5 milliseconds and the sliding distance is 3 milliseconds, then the first first audio segment is an audio segment of 0 to 5 milliseconds, the second first audio segment is an audio segment of 3 to 8 milliseconds, the third first audio segment is an audio segment of 6 to 11 milliseconds, and so on.

[0090] Step A22: Extract features from multiple first audio segments to obtain the features of each first audio segment.

[0091] The feature extractor can be invoked to extract temporal features from each first audio segment, obtaining the temporal features of each first audio segment, and using these temporal features as the features of each first audio segment. Alternatively, the feature extractor can be invoked to extract frequency-domain features from each first audio segment, obtaining the frequency-domain features of each first audio segment, and using these frequency-domain features as the features of each first audio segment. Alternatively, the temporal features of any first audio segment can be concatenated with its frequency-domain features, and the resulting concatenated features can be used as the features of that first audio segment.

[0092] Step A23: Based on the characteristics of each first audio segment, classify each first audio segment to obtain the unit beat type to which each first audio segment belongs.

[0093] A classifier can be invoked to classify the first audio segment based on its features, obtaining the probability that the first audio segment belongs to each beat type. Optionally, the beat types include strong beat, secondary strong beat, and weak beat. In the music field, a strong beat is the first beat of each measure, a secondary strong beat is a beat with a strong note, and a weak beat is a beat without a strong note.

[0094] The higher the probability that the first audio segment belongs to any beat type, the more likely it is to belong to that beat type. Based on this, we can select the beat type with the highest probability from the probabilities of the first audio segment belonging to each beat type, and take the beat type corresponding to the highest probability as the beat type to which the first audio segment belongs. In this way, we can determine the beat type to which each first audio segment belongs.

[0095] Step A24: Based on the beat type of each first audio segment, determine multiple audio transition points.

[0096] Based on the beat type of each first audio segment, multiple first audio segments are merged into multiple beat audio segments. Each beat audio segment includes at least one first audio segment. Each beat audio segment includes an audio segment of one beat, and any two beat audio segments correspond to the same type of beat or different types of beats. In music, a beat is a unit of time that reflects the combination rules of beat types. For example, a weak beat and a strong beat form one type of beat, and two weak beats and a strong beat form another type of beat.

[0097] Next, based on any two adjacent beat audio segments, an audio transition point is determined (this audio transition point is the first transition point mentioned below). Optionally, if there is no overlapping audio segment between any two adjacent beat audio segments, then any moment between the two adjacent beat audio segments is taken as the audio transition point. For example, if the first beat audio segment is from 0 to 15 milliseconds and the second beat audio segment is from 15 to 40 milliseconds, then 15 milliseconds is the audio transition point. Alternatively, if there is an overlapping audio segment between any two adjacent beat audio segments, then any moment within the overlapping audio segment is taken as the audio transition point. For example, if the first beat audio segment is from 0 to 20 milliseconds and the second beat audio segment is from 15 to 40 milliseconds, then 18 milliseconds is the audio transition point.

[0098] Optionally, if any two adjacent beat audio segments correspond to different types of beats, an audio transition point is determined based on these two adjacent beat audio segments; if any two adjacent beat audio segments correspond to the same type of beat, then it is not necessary to determine an audio transition point based on these two adjacent beat audio segments.

[0099] By using steps A21 to A24, beat detection processing can be performed on the audio information to obtain multiple beat audio segments. This allows for the determination of the playback duration of each beat in the background music, thereby enabling the identification of audio transition points between beat audio segments on the playback timeline and obtaining a playback timeline with each beat's duration. Determining a playback timeline with each beat's duration facilitates switching the virtual camera's movement information at beat transition points in the background music, improving the visual presentation of the scene.

[0100] In another possible implementation, the content detection process is a structure detection process. In this case, step A2 includes steps A25 to A28 (not shown in the figure).

[0101] Step A25 involves segmenting the audio information to obtain multiple second audio segments. The implementation of step A25 is similar to that described in step A21, and will not be repeated here. Optionally, any second audio segment may be a first audio segment, or it may not be a single first audio segment.

[0102] Step A26 involves extracting features from multiple second audio segments to obtain the features of each second audio segment. The implementation of step A26 is described in the following section, as the underlying principles are similar and will not be repeated here. Optionally, the features of any second audio segment may be features of a first audio segment, or they may not be features of any first audio segment.

[0103] Step A27: Based on the features of each second audio segment, cluster the second audio segments to obtain at least one cluster, where each cluster includes at least one second audio segment. For ease of description, each cluster here is denoted as a target cluster.

[0104] In one possible implementation, at least one initial cluster center can be obtained, where each initial cluster center represents a defined feature and corresponds to a specific cluster. For any second audio segment, the distances between the features of the second audio segment and each of the initial cluster centers can be calculated. The cluster corresponding to the initial cluster center with the smallest distance is then taken as the cluster to which the second audio segment belongs. In this way, each second audio segment can be clustered into the cluster corresponding to its respective initial cluster center.

[0105] If the clusters corresponding to each initial cluster center satisfy the clustering termination condition, then the clusters corresponding to each initial cluster center are designated as target clusters. This application does not limit the number of clusters corresponding to each initial cluster center to satisfying the clustering termination condition. For example, the clustering termination condition for each initial cluster center is satisfied when the number of clustering iterations for each initial cluster center reaches a set number. Alternatively, based on the features of each second audio segment included in the cluster corresponding to any initial cluster center, an average calculation is performed to obtain a first cluster center. The distance between the first cluster center and the initial cluster center is calculated; if the distance is less than a threshold, then the cluster corresponding to that initial cluster center satisfies the clustering termination condition.

[0106] If the clusters corresponding to the initial cluster centers do not meet the clustering termination condition, the features of each second audio segment included in the clusters corresponding to the initial cluster centers are averaged to obtain the first cluster centers. For any second audio segment, the distance between the features of the second audio segment and each of the first cluster centers can be calculated, and the cluster corresponding to the first cluster center with the smallest distance is taken as the cluster to which the second audio segment belongs. In this way, each second audio segment can be clustered into the cluster corresponding to each of the first cluster centers.

[0107] If the clusters corresponding to each of the first cluster centers satisfy the clustering termination condition, then the clusters corresponding to each of the first cluster centers are designated as target clusters. If the clusters corresponding to each of the first cluster centers do not satisfy the clustering termination condition, then the target clusters are determined based on the clusters corresponding to each of the first cluster centers. This process continues, and will not be elaborated further here.

[0108] In another possible implementation, for any two second audio segments, the distance between their features can be calculated. If the distance between the features of the two second audio segments is less than a threshold, then the two second audio segments are clustered in the same target cluster; if the distance between the features of the two second audio segments is not less than the threshold, then the two second audio segments are clustered in different target clusters. In this way, each second audio segment can be clustered into its respective target cluster.

[0109] By clustering the various second audio segments, similar second audio segments can be grouped into the same target cluster, making the second audio segments included in the same target cluster similar, while the second audio segments included in different target clusters are quite different.

[0110] In the field of music, background music can be categorized structurally into several audio formats, including intro, verse, chorus, interlude, and outro. The intro is the beginning of the music without lyrics; the verse is the prelude; the chorus is generally the climax; the interlude connects the previous chorus with the next verse and usually has no lyrics; and the outro is the end of the music without lyrics.

[0111] Typically, the audio structures of different musical forms differ significantly. Therefore, by clustering each second audio segment into a specific target cluster, the background music can be divided from a musical structure perspective, resulting in audio with at least one musical form. Each target cluster corresponds to an audio segment of a musical form, and any two target clusters may correspond to the same or different musical forms. For example, the intro, interlude, and outro are usually relatively soothing, serving as introductions and connections. Therefore, any musical form in the intro, interlude, or outro may result in a stylistic shift, leading to the audio of that musical form being clustered into at least two target clusters during the clustering process.

[0112] Step A28: Determine multiple audio inflection points based on at least one cluster. For ease of description, any cluster here will be referred to as a target cluster.

[0113] For any given target cluster, the various second audio segments included in this target cluster can be merged to obtain at least one structured audio segment, and any structured audio segment includes at least one second audio segment. For example, if a target cluster includes second audio segments of 0-30 seconds, 30-45 seconds, and 120-180 seconds, then these second audio segments can be merged into a structured audio segment of 0-45 seconds and a structured audio segment of 120-180 seconds.

[0114] Next, based on any two adjacent structural audio segments, an audio transition point is determined (this audio transition point is the second transition point mentioned below). Optionally, if there is no overlapping audio segment between any two adjacent structural audio segments, then any moment between the two adjacent structural audio segments is taken as the audio transition point. Alternatively, if there is an overlapping audio segment between any two adjacent structural audio segments, then any moment within the overlapping audio segment is taken as the audio transition point.

[0115] Understandably, when a target cluster corresponds to an audio segment with a musical structure, merging the various second audio segments included in any target cluster into at least one structural audio segment allows this structural audio segment to be an audio segment representing a musical structure within the background music. For example, the structural audio segment could be any of the musical structures such as intro, verse, chorus, interlude, or outro. Determining audio transition points through adjacent structural audio segments facilitates switching virtual camera movement information at these transition points in the background music's musical structure, enhancing the visual presentation of the scene.

[0116] By using steps A25 to A28, the audio information can be processed to detect the structure, resulting in multiple structured audio segments. This allows the playback duration of each musical structure in the background music to be determined, thereby enabling the determination of audio transition points between structured audio segments on the playback timeline and obtaining a playback timeline with the duration of each musical structure.

[0117] Step A3: Determine multiple target inflection points based on multiple audio inflection points.

[0118] For example, each audio transition point can be used as a target transition point. Alternatively, a portion of the multiple audio transition points can be used as target transition points, where there is a random number of audio transition points between two adjacent target transition points, or the playback duration between two adjacent target transition points is greater than a set duration.

[0119] Optionally, the audio transition points determined through steps A21 to A24 are recorded as first transition points, and the audio transition points determined through steps A25 to A28 are recorded as second transition points. When a first transition point and a second transition point coincide, or when the duration between a first transition point and a second transition point is less than a threshold, a target transition point is determined based on this first transition point and this second transition point. In other words, the beat transition points (i.e., first transition points) can be filtered based on the transition points of the musical structure (i.e., the second transition point), so that the target transition point is both a transition point of the musical structure and a beat transition point, making the switching of camera movement information more in line with the musical rhythm, which is beneficial to improving the performance of the scene.

[0120] In implementation B, the frame file includes a lyrics file. In this case, step 202 includes steps B1 to B3 (not shown in the figure).

[0121] Step B1: Extract the lyrics text from the background music file.

[0122] Since the background music file includes frame files, and the frame files include lyric files, the lyric files can be extracted from the background music file, and the lyric text contained within the lyric files can be obtained. Optionally, the lyric text includes multiple lyric phrases and the playback time occupied by each lyric phrase.

[0123] Step B2 involves performing sentiment analysis on the lyrics text to obtain multiple emotional turning points on the playback timeline.

[0124] The electronic device is equipped with a toolkit for sentiment analysis of song lyrics. By calling the toolkit to perform sentiment analysis on the song lyrics, multiple playback moments can be obtained, each of which represents an emotional turning point on the playback timeline. This application does not limit the toolkit; for example, the toolkit is a Chinese Sentiment (Cnsenti) library, which can be used to perform sentiment analysis on song lyrics in Chinese.

[0125] In an exemplary embodiment, step B2 includes steps B21 to B23 (not shown in the figures).

[0126] Step B21 involves classifying the lyrics text to obtain the global emotional information of the lyrics text and the local emotional information of each lyric phrase included in the lyrics text.

[0127] In this embodiment, language analysis can be performed on the lyrics text to obtain the language to which the lyrics text belongs (i.e., the language of the lyrics), for example, the language to which the lyrics text belongs is Chinese or Japanese. Next, an emotional lexicon corresponding to the language to which the lyrics text belongs is obtained. This emotional lexicon includes multiple emotional information, such as positive emotions, neutral emotions, and negative emotions.

[0128] The lyrics text can be input into a classifier, which determines the probability that the lyrics text belongs to each sentiment information included in the sentiment lexicon. Since the higher the probability that the lyrics text belongs to any sentiment information, the more likely the lyrics text is to belong to that sentiment information, the highest probability can be selected from the probabilities of the lyrics text belonging to each sentiment information, and the sentiment information corresponding to the highest probability can be taken as the global sentiment information to which the lyrics text belongs.

[0129] The lyrics text can be divided into individual lyric phrases. Each lyric phrase is then input into a classifier, which determines the probability that it belongs to any of the various sentiment information categories included in the sentiment lexicon. Since a higher probability that a lyric phrase belongs to any particular sentiment information means that phrase is more likely to belong to that sentiment, the lyric phrase with the highest probability can be selected as the local sentiment information to which it belongs. In this way, the local sentiment information to which each lyric phrase belongs can be determined.

[0130] Step B22: For any lyric phrase, based on the global sentiment information and the local sentiment information to which the lyric phrase belongs, determine the sentiment-related index corresponding to the lyric phrase. The sentiment-related index corresponding to the lyric phrase is used to characterize the degree of correlation between the local sentiment information to which the lyric phrase belongs and the global sentiment information.

[0131] The emotional correlation index between every two emotional information items included in the emotional lexicon can be configured, so as to determine the emotional correlation index between the global emotional information to which the lyrics text belongs and the local emotional information to which each lyric sentence belongs based on the configuration information.

[0132] Optionally, the more relevant the local emotional information of any lyric phrase is to the global emotional information of the lyric text, the greater the emotional correlation index between the global emotional information of the lyric text and the local emotional information of that lyric phrase; conversely, the less relevant the local emotional information of any lyric phrase is to the global emotional information of the lyric text, the smaller the emotional correlation index between them. Therefore, the emotional correlation index between the global emotional information of the lyric text and the local emotional information of any lyric phrase can characterize the degree of correlation between them.

[0133] This application does not limit the value of the sentiment-related index; the sentiment-related index can be any data, such as -3 or 10. In an exemplary embodiment, the sentiment-related index is a data greater than or equal to 0 and less than or equal to 1. In this case, the sentiment-related index can be recorded as a weight.

[0134] Step B23: Based on the emotional relevance indicators corresponding to each lyric, determine multiple emotional turning points.

[0135] For any two adjacent lyrics, if the difference between the emotional correlation indicators corresponding to these two lyrics is greater than a set threshold, it is determined that there is an emotional turning point between the playback time occupied by these two lyrics; if the difference between the emotional correlation indicators corresponding to these two lyrics is not greater than the set threshold, it is determined that there is no emotional turning point between the playback time occupied by these two lyrics.

[0136] For example, if the sentiment correlation index for lyric line A is 0.9 and the sentiment correlation index for lyric line B is 0.2, and 0.9 - 0.2 = 0.7 > 0.65 (the set threshold), then an sentiment inflection point is determined between the playback time occupied by lyric line A and the playback time occupied by lyric line B. Since lyric line A occupies 0 to 5 seconds of playback time and lyric line B occupies 5 to 9 seconds of playback time, 5 seconds is considered a sentiment inflection point.

[0137] By using steps B21 to B23, sentiment analysis can be performed on the lyrics text to obtain the correlation between the local and global sentiment information of each lyric phrase. Based on this correlation, each emotional turning point can be determined, which is beneficial for switching the camera movement information of the virtual camera at the emotional turning points of the background music and improving the performance of the scene.

[0138] Step B3: Identify multiple target turning points based on multiple emotional turning points.

[0139] For example, each emotional turning point can be used as a target turning point. Alternatively, some emotional turning points from a set of emotional turning points can be used as target turning points, with a random number of emotional turning points between adjacent target turning points, or the playback duration between adjacent target turning points exceeding a set duration. It is understood that camera movement information is typically not switched within a complete lyric phrase; therefore, no single target turning point will fall within the playback time occupied by a single lyric phrase.

[0140] In implementation C, the frame file includes an audio file and a lyrics file. In this case, step 202 includes steps C1 to C4 (not shown in the figure).

[0141] Step C1 involves extracting audio information and lyrics from the background music file. The description of step C1 can be found in steps A1 and B1, which are similar in principle and will not be repeated here.

[0142] Step C2 involves performing content detection processing on the audio information to obtain multiple audio inflection points on the playback timeline. The description of step C2 is the same as that of step A2, as their implementation principles are similar and will not be repeated here.

[0143] Step C3 involves performing sentiment analysis on the lyrics to obtain multiple emotional turning points along the playback timeline. The description of step C3 is similar to that of step B2, and therefore will not be repeated here.

[0144] Step C4 involves identifying multiple target turning points based on multiple audio and emotional turning points. The description of step C4 can be found in steps A3 and B3, as their implementation principles are similar and will not be repeated here.

[0145] Optionally, when a first turning point, a second turning point, and an emotional turning point coincide, or when the maximum duration between a first turning point, a second turning point, and an emotional turning point is less than a threshold, a target turning point is determined based on this first turning point, this second turning point, and this emotional turning point. In other words, beat turning points (i.e., first turning points) can be filtered based on turning points in the musical structure (i.e., second turning points) and turning points in emotional information (i.e., emotional turning points), so that the target turning point is a turning point in the musical structure, a beat turning point, and an emotional turning point. This makes the switching of camera movement information more consistent with the musical rhythm and the expression of emotion, which is beneficial to improving the expressive effect of the scene.

[0146] It is understandable that, besides parsing and processing background music files to obtain the target turning point, there can be other ways to determine the target turning point. For example, when there are virtual objects in a virtual scene, and these virtual objects are in motion, the motion information of the virtual objects (such as dance videos) can be obtained. By parsing the motion information of the virtual objects, multiple motion turning points of the virtual objects during the motion process can be obtained. Based on at least one of the multiple motion turning points, multiple audio turning points, and multiple emotional turning points, the target turning point can be determined.

[0147] Step 203: Determine multiple camera movement switching moments on the playback timeline based on multiple target turning points.

[0148] For example, each target turning point is used as a camera movement transition moment. Alternatively, some target turning points from a plurality of target turning points are used as camera movement transition moments, where a random number of target turning points exist between two adjacent camera movement transition moments, or where the playback duration between two adjacent camera movement transition moments is greater than a set duration. Since camera movement information is typically not switched within a complete lyric phrase, any camera movement transition moment does not fall within the playback time occupied by a lyric phrase.

[0149] Step 204: For any camera movement switch moment, select the target camera movement information from multiple candidate camera movement information.

[0150] The electronic device allows users to switch camera movement information at any camera movement transition, or not to switch it. When switching camera movement information, users can select the target camera movement information from multiple candidate camera movement information options.

[0151] Optionally, in step 204, “selecting target camera movement information from multiple candidate camera movement information” includes step 2041 or step 2042 (not shown in the figure).

[0152] Step 2041: Randomly generate target camera movement number, select the target camera movement information corresponding to the target camera movement number from multiple candidate camera movement information, and each candidate camera movement information corresponds to a camera movement number.

[0153] In this embodiment, each candidate camera movement information corresponds to a unique camera movement sequence number. The electronic device can randomly generate a target camera movement sequence number, which is one of multiple camera movement sequence numbers. The electronic device selects the candidate camera movement information corresponding to the target camera movement sequence number as the target camera movement information.

[0154] Step 2042: Based on the operation of selecting any candidate camera movement information in the camera movement display list, determine any candidate camera movement information as the target camera movement information. The camera movement display list is used to display multiple candidate camera movement information.

[0155] In this embodiment, the electronic device can display a camera movement display list on a screen, which shows multiple candidate camera movement information. An object can select one candidate camera movement information from the list, and the electronic device, in response to this selection, uses the selected candidate camera movement information as the target camera movement information.

[0156] Understandably, since each candidate shot information corresponds to a shot sequence number, electronic devices can display each shot sequence number in the shot display list, thereby displaying each candidate shot information.

[0157] Please see Figure 3 , Figure 3 This is a schematic diagram of a display page provided in an embodiment of this application. The display page is roughly divided into four display areas, which will be described in detail below.

[0158] The first display area shows a virtual scene and at least one virtual camera (in) Figure 3 It shows two virtual cameras and a functional area 2. Each virtual camera corresponds to a candidate camera movement information, used to capture a virtual scene based on the candidate camera movement information corresponding to that virtual camera. Functional area 2 displays at least one functional control. This embodiment does not limit the functional controls displayed in functional area 2. Figure 3The functional controls shown are merely illustrative. Among them, Figure 3 The functional controls shown in the functional area 2 are, in order: a control for opening a document, a control for saving a document, a control for taking a picture, a control for viewing information, a control for modifying information, a control for swiping left, and a control for swiping right.

[0159] The second display area shows the current camera movement transition time and the camera movement display list. The "Current Time: 90" indicates that the current camera movement transition time is the playback time corresponding to 90 milliseconds (ms). The camera movement display list shows multiple camera movement sequences, each sequence corresponding to a candidate camera movement; therefore, the camera movement display list effectively displays multiple candidate camera movement information. Figure 3 The camera movement display list includes: camera movement sequence 1, camera movement sequence 16, camera movement sequence 13, camera movement sequence 12, camera movement sequence 11, camera movement sequence 10, and camera movement sequence 6.

[0160] The third display area includes a functional area 1, a "+Segment" control for adding screen clips, a search box, and at least one area for displaying screen clips. Functional area 1 displays at least one functional control; however, this embodiment does not limit the functional controls displayed in functional area 1. Figure 3 The functional controls shown are merely illustrative. The search box supports information input; by entering search terms in the search box, content matching the search terms can be found in the database. The area for displaying screen clips includes the displayed screen clips and functional area 3. Functional area 3 displays at least one functional control; this embodiment does not limit the functional controls displayed in functional area 3. Figure 3 The functional controls shown are merely illustrative. Figure 3 In the middle, the function controls displayed in function area 3 are, in order: the function control "+" for adding screen clips and the function control for capturing screen footage.

[0161] The fourth display area shows the playback timeline, an area for displaying the scene footage captured by the virtual camera based on at least one target camera movement information, an area for displaying the scene footage captured by the virtual camera based on any candidate camera movement information, a slider for vertical scrolling, and a slider for horizontal scrolling.

[0162] The playback timeline displays multiple playback moments and an indicator (i.e., a black arrow) to indicate the current moment. Some of the playback moments are camera transition moments. For example, the playback moments corresponding to 45 milliseconds, 90 milliseconds, and 135 milliseconds are camera transition moments, and the current moment is the playback moment corresponding to 90 milliseconds.

[0163] The areas used to display scene images obtained by a virtual camera based on at least one target movement information are: scene images obtained by a virtual camera based on the target movement information corresponding to movement number 1, scene images obtained by a virtual camera based on the target movement information corresponding to movement number 6, scene images obtained by a virtual camera based on the target movement information corresponding to movement number 6, and scene images obtained by a virtual camera based on the target movement information corresponding to movement number 13.

[0164] The areas used to display the scene images obtained by the virtual camera based on any candidate camera movement information are: the scene image obtained by the virtual camera based on the target camera movement information corresponding to camera movement number 1 (this scene image is a panorama), and the scene image obtained by the virtual camera based on the target camera movement information corresponding to camera movement number 16 (this scene image is a close-up).

[0165] In photography, with a fixed camera focal length, the subject appears in varying sizes within the camera's view due to differences in the distance between the camera and the subject. Generally, based on the size of the subject within the camera's view, five shot types can be distinguished: close-up, medium shot, long shot, and extreme shot. Taking the human body as an example, a close-up corresponds to the area above the shoulders, a medium shot to the chest, a long shot to the knees, a full shot to the entire body and surrounding environment, and an extreme shot to the entire environment surrounding the body.

[0166] In this embodiment of the application, the electronic device can determine the target camera information corresponding to any camera movement switching moment in accordance with the above steps 2041 or 2042. The target camera movement information is used to enable the virtual camera to start shooting the virtual scene from any camera movement switching moment based on the target camera movement information.

[0167] Optionally, the electronic device can randomly generate target camera movement numbers corresponding to each camera movement switching moment to select the target camera movement information corresponding to the target camera movement number at each camera movement switching moment. If the object does not meet the target camera movement information (i.e., the original camera movement information mentioned below) for a certain camera movement switching moment, the electronic device reads multiple candidate camera movement information and displays the camera movement numbers of each candidate camera movement information in the camera movement display list. Based on the object's operation of selecting any camera movement number in the camera movement display list, the electronic device obtains the candidate camera movement information corresponding to the selected camera movement number and generates a switching control. Based on the object's confirmation operation of the switching control, the electronic device reads the currently targeted camera movement switching moment on the playback timeline and switches the original camera movement information to the candidate camera movement information corresponding to the selected camera movement number.

[0168] By determining the specific moments of each camera movement transition and randomly selecting the target camera movement information corresponding to each transition moment, it is possible to automatically generate camera movement transition moments that match the rhythm, emotional expression, and musical structure of background music in virtual scenes accompanied by arbitrary audio, music, or songs, for situations such as virtual human performances, live broadcasts, or computer graphics (CG) animations. Furthermore, a solution is provided for manually adjusting the target camera movement information corresponding to any camera movement transition moment, supporting rapid modification of camera movement information, saving manpower and time costs, and resulting in better music-integrated effects.

[0169] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the background music text and candidate camera movement information involved in this application were obtained with full authorization.

[0170] The method described above parses the background music file to obtain multiple target turning points on the playback timeline. Based on each target turning point, it determines multiple camera movement transition moments on the playback timeline and identifies the target camera movement information corresponding to each transition moment. By determining the target camera movement information corresponding to each transition moment, it enables switching the virtual camera's movement information at the turning points of the background music, thereby improving the performance of the scene footage captured by the virtual camera.

[0171] This application provides a method for displaying scene images. This method can be applied to the above-described implementation environment and can switch the camera movement information of the virtual camera, thereby improving the performance of scene images obtained by the virtual camera shooting virtual scenes. Figure 4 The flowchart shown in this embodiment of the application illustrates a method for displaying a scene image, in which the terminal device 101 executes the method for displaying a scene image according to this embodiment. Figure 4 As shown, the method includes the following steps.

[0172] Step 401: Display the live streaming interface and play background music. The live streaming interface includes a first scene image, which is a virtual scene image obtained by shooting the virtual scene with a virtual camera based on the first camera movement information.

[0173] A virtual camera can capture a virtual scene based on first camera movement information to obtain a first scene image. The first camera movement information is the target camera movement information corresponding to any camera movement switch moment in step 204. The virtual camera can start capturing the virtual scene from that camera movement switch moment based on the target camera movement information corresponding to that switch moment. The first scene image is rendered by a server or terminal device, allowing it to be displayed on the terminal device's live streaming interface, while simultaneously playing the background music from step 201. The aforementioned technology for live streaming scene images obtained from capturing a virtual scene is also called virtual live streaming technology.

[0174] Virtual scenes can be constructed based on 3D modeling technology and include at least one of static object models and dynamic object models. In an exemplary embodiment, the virtual scene includes a virtual human, and virtual live streaming is the live streaming of a virtual human. The virtual human is constructed using 3D modeling technology, and its movement is driven by the motion capture data of the virtual human. The virtual human is a dynamic object model. The virtual human can also be called a motion capture actor. By collecting data such as facial expressions and body movements of the virtual human, motion capture data is obtained, and the virtual human's movement is driven based on this motion capture data, making the virtual human's movement synchronized with the virtual human's movement. A virtual camera captures the scene of the virtual human in motion, and a game engine renders the scene, allowing the live streaming interface to display the scene of the virtual human's movement.

[0175] Optionally, the virtual scene can also include the virtual environment in which the virtual person resides. A virtual environment can be constructed using 3D modeling technology, allowing the virtual person to move within it. By using a virtual camera to film the virtual person in motion within the virtual environment, the scene of the virtual person moving within the virtual environment can be displayed on a live streaming interface.

[0176] Step 402: In response to the background music playing at the moment of camera movement switching, switch the first camera movement information to the second camera movement information.

[0177] The first camera movement information corresponds to a camera movement transition moment. When the background music plays to the next camera movement transition moment corresponding to the first camera movement information, the first camera movement information can be switched to the target camera movement information corresponding to the next camera movement transition moment, and this target camera movement information can be recorded as the second camera movement information. It can be understood that any camera movement transition moment is a moment on the background music playback timeline, and this camera movement transition moment corresponds to a target turning point in the background music.

[0178] In one possible implementation, step 402 includes: displaying a camera movement switching control on the live streaming interface; and switching the first camera movement information to the second camera movement information based on a confirmation operation on the camera movement switching control.

[0179] When the background music reaches the moment of camera movement change, a camera movement change control is displayed on the live stream interface. If the viewer confirms the camera movement change control (for example, if the viewer clicks the camera movement change control within a set time), the first camera movement information is switched to the second camera movement information based on the confirmation.

[0180] Step 403: Display the second scene on the live streaming interface. The second scene is a picture of a virtual scene captured by a virtual camera based on the second camera movement information.

[0181] The virtual camera captures the virtual scene based on the second camera movement information, and the resulting scene image is recorded as the second scene image. The second scene image is rendered by the game engine, which can be displayed in the live broadcast interface and background music is played synchronously.

[0182] Understandably, when the camera movement switching control is displayed on the live stream interface, any viewer can confirm or cancel the camera movement switching control (for example, if a viewer does not click the camera movement switching control within a set time). If a viewer confirms the camera movement, the first camera movement information can be switched to the second camera movement information; if a viewer cancels the camera movement, the first camera movement information does not need to be switched.

[0183] In an exemplary embodiment, each camera movement transition moment corresponds to at least one target camera movement information. In response to background music playing to a camera movement transition moment, camera movement transition controls corresponding to each target camera movement information are displayed on the live streaming interface. Based on a confirmation operation for any camera movement transition control, the target camera movement information corresponding to that control is determined as the second camera movement information, and the first camera movement information is switched to the second camera movement information; based on a cancellation operation for each camera movement transition control, the first camera movement information is not switched.

[0184] The server can render both the first and second scene images. When a viewer confirms a live stream, the server sends the second scene image to the terminal device based on the confirmation information, allowing the terminal device to display the second scene image. Conversely, when a viewer cancels a live stream, the server sends the first scene image to the terminal device based on the cancellation information, allowing the terminal device to display the first scene image.

[0185] In this embodiment, based on the confirmation or cancellation operation of the live stream viewer on the camera movement switching control, the corresponding scene screen is pushed to the live stream viewer. This not only realizes the switching of the virtual camera movement at the turning point of the background music to improve the shooting effect, but also meets the viewing needs of different live stream viewers and improves the live stream effect.

[0186] It is understandable that the methods for determining the timing of camera movement transitions and the methods for determining the target turning point in steps 401 to 403 can be found in... Figure 2 The methods for determining relevant camera movement information will not be elaborated here.

[0187] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the first camera movement information and the second camera movement information involved in this application were obtained with full authorization.

[0188] The above method involves switching the camera movement information of the virtual camera during the display of the live broadcast interface and the playback of background music, when the background music reaches the camera movement switching moment. The live broadcast interface then displays the scene footage captured based on the switched camera movement information. By switching the camera movement information of the virtual camera, the performance effect of the scene footage captured by the virtual camera on the virtual scene is improved.

[0189] The above describes the method for determining camera movement information provided in the embodiments of this application from the perspective of method steps. The following section, in conjunction with... Figure 5 and Figure 6 To describe further. Please see [link / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating the determination of camera movement information provided in an embodiment of this application.

[0190] In this embodiment, camera movement information is configured. A subject can be selected, which is the virtual scene mentioned above. The number of virtual cameras is configured, with one virtual camera corresponding to one candidate camera movement information. The parameters and motion trajectories of each virtual camera can be configured to achieve the camera movement information configuration. A candidate camera movement information includes the parameters and motion trajectory of one virtual camera. On the other hand, the background music file can be parsed to obtain multiple camera movement transition moments, then a playback timeline is generated, and each camera movement transition moment is marked on the playback timeline. Afterwards, the camera movement information corresponding to each camera movement transition moment is determined.

[0191] The process of parsing the background music file to obtain multiple camera movement transition moments can be found in [reference needed]. Figure 6 , Figure 6 This is a schematic diagram illustrating the determination of camera movement switching moments according to an embodiment of this application. In this embodiment, a background music file can be input into an electronic device, and the background music file can be parsed to obtain a file header, a frame file, and a file footer. The frame file includes at least one of audio information and lyrics text.

[0192] In one possible implementation, the electronic device loads audio information. On one hand, it performs beat detection on the audio information to obtain the individual beats of the background music, determines the duration of each beat, and identifies multiple audio transition points based on the duration of each beat. On the other hand, it performs musical structure detection on the audio information to obtain the various musical structures, determines the duration of each musical structure, and identifies multiple audio transition points based on the duration of each musical structure.

[0193] In another possible implementation, the electronic device loads the lyrics text and determines the language of the lyrics. Then, based on a sentiment analysis library corresponding to the language of the lyrics, sentiment analysis is performed on the lyrics text to obtain the sentiment analysis results. The sentiment analysis results include global sentiment information of the lyrics text and local sentiment information of each lyric phrase within the lyrics text. Based on the sentiment analysis results, the weight of each lyric phrase is determined; the weight of any lyric phrase is the sentiment-related index mentioned above. Subsequently, multiple sentiment inflection points are determined based on the weights of each lyric phrase.

[0194] Next, based on multiple audio turning points and multiple emotional turning points, multiple camera movement switching moments are determined to identify the camera movement information corresponding to each camera movement switching moment.

[0195] In this embodiment, multiple audio transition points are obtained by performing beat detection and musical structure detection on the audio information, and multiple emotional transition points are obtained by performing sentiment analysis on the lyrics. Based on each audio transition point and each emotional transition point, multiple camera movement switching moments are determined, and the target camera movement information corresponding to each camera movement switching moment is also determined. By determining the target camera movement information corresponding to the camera movement switching moments, the camera movement information of the virtual camera is switched at the beat transition points, musical structure transition points, and emotional transition points of the background music, thereby improving the performance effect of the scene images obtained by the virtual camera shooting the virtual scene.

[0196] Figure 7 The diagram shown is a structural schematic of a device for determining camera movement information provided in an embodiment of this application. Figure 7 As shown, the device includes:

[0197] The acquisition module 701 is used to acquire the background music file. The background music is the music that accompanies the display of the scene. The scene is the picture obtained by capturing the virtual scene with a virtual camera.

[0198] The acquisition module 701 is also used to acquire multiple candidate camera movement information, where any candidate camera movement information is information from a virtual camera.

[0199] The parsing module 702 is used to parse and process the background music file to obtain multiple target turning points on the playback timeline of the background music;

[0200] The determination module 703 is used to determine multiple camera switching moments on the playback timeline based on multiple target turning points;

[0201] Selection module 704 is used to select target camera information from multiple candidate camera information for any camera movement switching moment. The target camera information is used to enable the virtual camera to start shooting the virtual scene from any camera movement switching moment based on the target camera information.

[0202] In one possible implementation, the acquisition module 701 is used to acquire the position information of multiple trajectory points on the motion trajectory of the virtual camera for any candidate camera movement information; determine the motion trajectory of the virtual camera based on the position information of the multiple trajectory points; set the motion information of a reference object on the motion trajectory to obtain any candidate camera movement information, and the reference object is used to make the virtual camera follow the reference object and move on the motion trajectory according to the motion information of the reference object.

[0203] In one possible implementation, the parsing module 702 is used to extract audio information from the background music file; perform content detection processing on the audio information to obtain multiple audio turning points on the playback timeline; and determine multiple target turning points based on the multiple audio turning points.

[0204] In one possible implementation, the parsing module 702 is used to segment the audio information to obtain multiple first audio segments; extract features from the multiple first audio segments to obtain the features of each first audio segment; classify each first audio segment based on the features of each first audio segment to obtain the beat type to which each first audio segment belongs; and determine multiple audio inflection points based on the beat type to which each first audio segment belongs.

[0205] In one possible implementation, the parsing module 702 is used to segment the audio information to obtain multiple second audio segments; extract features from the multiple second audio segments to obtain features of each second audio segment; perform clustering processing on each second audio segment based on the features of each second audio segment to obtain at least one cluster, where each cluster includes at least one second audio segment; and determine multiple audio inflection points based on at least one cluster.

[0206] In one possible implementation, the parsing module 702 is used to extract lyrics text from the background music file; perform sentiment analysis on the lyrics text to obtain multiple emotional turning points on the playback timeline; and determine multiple target turning points based on the multiple emotional turning points.

[0207] In one possible implementation, the parsing module 702 is used to classify the lyrics text to obtain the global emotional information to which the lyrics text belongs and the local emotional information to which each lyric phrase belongs. For any lyric phrase, based on the global emotional information and the local emotional information to which the lyric phrase belongs, an emotional correlation index corresponding to the lyric phrase is determined. The emotional correlation index corresponding to the lyric phrase is used to characterize the degree of correlation between the local emotional information to which the lyric phrase belongs and the global emotional information. Based on the emotional correlation index corresponding to each lyric phrase, multiple emotional turning points are determined.

[0208] In one possible implementation, the selection module 704 is used to randomly generate a target camera movement number, select the target camera movement information corresponding to the target camera movement number from multiple candidate camera movement information, and each candidate camera movement information corresponds to a camera movement number; or, based on the operation of selecting any candidate camera movement information in the camera movement display list, any candidate camera movement information is determined as the target camera movement information, and the camera movement display list is used to display multiple candidate camera movement information.

[0209] The aforementioned device parses the background music file to obtain multiple target turning points on the playback timeline. Based on each target turning point, it determines multiple camera movement transition moments on the playback timeline and identifies the target camera movement information corresponding to each transition moment. By determining the target camera movement information corresponding to each transition moment, the device enables switching the virtual camera's movement information at the turning points of the background music, thereby improving the performance of the scene footage captured by the virtual camera.

[0210] It should be understood that the above Figure 7 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.

[0211] Figure 8 The diagram shown is a structural schematic of a scene display device provided in an embodiment of this application. Figure 8 As shown, the device includes:

[0212] Display module 801 is used to display the live broadcast interface, which includes a first scene image. The first scene image is a picture obtained by a virtual camera capturing a virtual scene based on first camera movement information.

[0213] Playback module 802 is used to play background music;

[0214] The switching module 803 is used to switch the first camera movement information to the second camera movement information in response to the background music playing at the camera movement switching moment.

[0215] The display module 801 is also used to display a second scene on the live streaming interface. The second scene is a picture of a virtual scene captured by a virtual camera based on the second camera movement information.

[0216] In one possible implementation, the switching module 803 is used to display a camera movement switching control on the live streaming interface; based on the confirmation operation of the camera movement switching control, the first camera movement information is switched to the second camera movement information.

[0217] The aforementioned device, while displaying the live streaming interface and playing background music, switches the camera movement information of the virtual camera when the background music reaches the camera movement switching moment. The live streaming interface then displays the scene footage captured based on the switched camera movement information. By switching the camera movement information of the virtual camera, the performance effect of the scene footage captured by the virtual camera on the virtual scene is improved.

[0218] It should be understood that the above Figure 9 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.

[0219] Figure 9 A structural block diagram of a terminal device 900 provided in an exemplary embodiment of this application is shown. The terminal device 900 includes a processor 901 and a memory 902.

[0220] Processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0221] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 are used to store at least one computer program, which is executed by the processor 901 to implement the method for determining camera movement information or the method for displaying scene images provided in the method embodiments of this application.

[0222] In some embodiments, the terminal device 900 may optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, and a power supply 908.

[0223] Peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 901 and memory 902. In some embodiments, processor 901, memory 902 and peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 901, memory 902 and peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0224] The radio frequency (RF) circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 904 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0225] Display screen 905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, disposed on the front panel of terminal device 900; in other embodiments, there may be at least two display screens 905, disposed on different surfaces of terminal device 900 or in a folded design; in other embodiments, display screen 905 may be a flexible display screen, disposed on a curved or folded surface of terminal device 900. Furthermore, display screen 905 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 905 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0226] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0227] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 901 for processing, or to the radio frequency circuit 904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal device 900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 907 may also include a headphone jack.

[0228] Power supply 908 is used to supply power to the various components in terminal device 900. Power supply 908 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 908 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0229] In some embodiments, the terminal device 900 further includes one or more sensors 909. The one or more sensors 909 include, but are not limited to, an accelerometer 911, a gyroscope 912, a pressure sensor 913, an optical sensor 914, and a proximity sensor 915.

[0230] Accelerometer 911 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 900. For example, accelerometer 911 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 901 can control display screen 905 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 911. Accelerometer 911 can also be used for games or for acquiring user motion data.

[0231] The gyroscope sensor 912 can detect the orientation and rotation angle of the terminal device 900. The gyroscope sensor 912, in conjunction with the accelerometer sensor 911, can collect the user's 3D movements on the terminal device 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0232] The pressure sensor 913 can be disposed on the side bezel of the terminal device 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side bezel of the terminal device 900, it can detect the user's grip signal on the terminal device 900, and the processor 901 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0233] An optical sensor 914 is used to collect ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 based on the ambient light intensity collected by the optical sensor 914. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera assembly 906 based on the ambient light intensity collected by the optical sensor 914.

[0234] The proximity sensor 915, also known as a distance sensor, is typically located on the front panel of the terminal device 900. The proximity sensor 915 is used to detect the distance between the user and the front of the terminal device 900. In one embodiment, when the proximity sensor 915 detects that the distance between the user and the front of the terminal device 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from a screen-on state to a screen-off state; when the proximity sensor 915 detects that the distance between the user and the front of the terminal device 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from a screen-off state to a screen-on state.

[0235] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the terminal device 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0236] Figure 10This is a schematic diagram of the server structure provided in the embodiments of this application. The server 1000 can vary considerably due to different configurations or performance. It may include one or more processors 1001 and one or more memories 1002. The one or more memories 1002 store at least one computer program, which is loaded and executed by the one or more processors 1001 to implement the method for determining camera movement information or the method for displaying scene images provided in the various method embodiments described above. For example, the processor 1001 is a CPU. Of course, the server 1000 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 may also include other components for implementing device functions, which will not be elaborated here.

[0237] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program, which is loaded and executed by a processor to enable an electronic device to implement any of the above-described methods for determining camera movement information or displaying scene images.

[0238] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0239] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer program, which is loaded and executed by a processor to enable an electronic device to implement any of the above-described methods for determining camera movement information or displaying scene images.

[0240] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0241] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0242] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for determining camera movement information, characterized in that, The method includes: Obtain the background music file, which is the music that accompanies the display of scene images, which are images obtained by capturing virtual scenes with a virtual camera; Acquire multiple candidate camera movement information, where any candidate camera movement information is information from the virtual camera. Extract the lyrics text from the background music file; The lyrics text is classified to obtain the global emotional information of the lyrics text and the local emotional information of each lyric phrase included in the lyrics text; For any lyric phrase, based on the global sentiment information and the local sentiment information to which the lyric phrase belongs, a sentiment-related index corresponding to the lyric phrase is determined. The sentiment-related index corresponding to the lyric phrase is used to characterize the degree of correlation between the local sentiment information to which the lyric phrase belongs and the global sentiment information. For each of the lyric phrases, if the difference between the emotional related indicators corresponding to the two adjacent lyric phrases is greater than a set threshold, it is determined that there is an emotional turning point between the playback time occupied by the two adjacent lyric phrases. Multiple target turning points on the playback timeline of the background music are determined based on multiple emotional turning points; Based on the multiple target turning points, multiple camera switching moments on the playback timeline are determined; For any camera movement switch moment, a target camera movement information is selected from the plurality of candidate camera movement information. The target camera movement information is used to enable the virtual camera to start shooting the virtual scene from the any camera movement switch moment based on the target camera movement information.

2. The method according to claim 1, characterized in that, The acquisition of multiple candidate camera movement information includes: For any candidate camera movement information, obtain the position information of multiple trajectory points on the motion trajectory of the virtual camera; The motion trajectory of the virtual camera is determined based on the position information of the multiple trajectory points; The motion information of a reference object is set on the motion trajectory to obtain the motion information of any candidate camera movement. The reference object is used to make the virtual camera follow the reference object and move on the motion trajectory according to the motion information of the reference object.

3. The method according to claim 1, characterized in that, The determination of multiple target turning points on the playback timeline of the background music based on multiple emotional turning points includes: Extract audio information from the background music file; The audio information is subjected to content detection processing to obtain multiple audio inflection points on the playback timeline; The multiple target turning points are determined based on the multiple audio turning points and the multiple emotional turning points.

4. The method according to claim 3, characterized in that, The content detection processing of the audio information to obtain multiple audio inflection points on the playback timeline includes: The audio information is segmented to obtain multiple first audio segments; Feature extraction is performed on the plurality of first audio segments to obtain the features of each first audio segment; Based on the characteristics of each first audio segment, the first audio segments are classified to obtain the beat type to which each first audio segment belongs; Based on the beat type to which each of the first audio segments belongs, the multiple audio transition points are determined.

5. The method according to claim 3, characterized in that, The content detection processing of the audio information to obtain multiple audio inflection points on the playback timeline includes: The audio information is segmented to obtain multiple second audio segments; Feature extraction is performed on the multiple second audio segments to obtain the features of each second audio segment; Based on the characteristics of each second audio segment, the second audio segments are clustered to obtain at least one cluster, and each cluster includes at least one second audio segment. The plurality of audio inflection points are determined based on the at least one cluster.

6. The method according to claim 1, characterized in that, The step of selecting the target camera movement information from the plurality of candidate camera movement information includes: Randomly generate a target camera movement sequence number, and select the target camera movement information corresponding to the target camera movement sequence number from the plurality of candidate camera movement information. Each candidate camera movement information corresponds to a camera movement sequence number. Alternatively, based on the operation of selecting any candidate camera movement information from the camera movement display list, the candidate camera movement information is determined as the target camera movement information, and the camera movement display list is used to display the multiple candidate camera movement information.

7. A method for displaying a scene, characterized in that, The method includes: The live stream interface is displayed and background music is played. The live stream interface includes a first scene screen, which is a screen obtained by shooting a virtual scene with a virtual camera based on first camera movement information. In response to the background music playing at the moment of camera movement switching, the first camera movement information is switched to the second camera movement information; The live streaming interface displays a second scene, which is a shot of the virtual scene captured by the virtual camera based on the second camera movement information. Wherein, the camera movement switching time, the first camera movement information, and the second camera movement information are all determined by the method according to any one of claims 1 to 6.

8. The method according to claim 7, characterized in that, The step of switching the first camera movement information to the second camera movement information includes: The live streaming interface displays camera movement switching controls; Based on the confirmation operation of the camera movement switching control, the first camera movement information is switched to the second camera movement information.

9. A device for determining camera movement information, characterized in that, The device includes: The acquisition module is used to acquire background music files. The background music is the music that accompanies the display of scene images, which are images obtained by capturing virtual scenes with a virtual camera. The acquisition module is also used to acquire multiple candidate camera movement information, where any one of the candidate camera movement information is information of the virtual camera. The parsing module is used to extract lyrics from the background music file; classify the lyrics to obtain global emotional information and local emotional information of each lyric phrase; for any lyric phrase, based on the global emotional information and the local emotional information, determine the corresponding emotional correlation index, which characterizes the correlation between the local emotional information and the global emotional information; for two adjacent lyric phrases, if the difference between the corresponding emotional correlation indices is greater than a set threshold, determine that there is an emotional turning point between the playback time occupied by the two adjacent lyric phrases; and determine multiple target turning points on the playback timeline of the background music based on multiple emotional turning points. The determination module is used to determine multiple camera movement switching moments on the playback timeline based on the multiple target turning points; The selection module is used to select target camera information from the plurality of candidate camera information at any camera movement switching moment. The target camera information is used to enable the virtual camera to start shooting the virtual scene from the any camera movement switching moment based on the target camera information.

10. The apparatus according to claim 9, characterized in that, The acquisition module is used to acquire the position information of multiple trajectory points on the motion trajectory of the virtual camera for any candidate camera movement information; determine the motion trajectory of the virtual camera based on the position information of the multiple trajectory points; set the motion information of a reference object on the motion trajectory to obtain the candidate camera movement information, wherein the reference object is used to make the virtual camera follow the reference object and move on the motion trajectory according to the motion information of the reference object.

11. The apparatus according to claim 9, characterized in that, The parsing module is used to extract audio information from the background music file; perform content detection processing on the audio information to obtain multiple audio turning points on the playback timeline; The multiple target turning points are determined based on the multiple audio turning points and the multiple emotional turning points.

12. The apparatus according to claim 11, characterized in that, The parsing module is used to segment the audio information to obtain multiple first audio segments; and to extract features from the multiple first audio segments to obtain the features of each first audio segment. Based on the characteristics of each first audio segment, the first audio segments are classified to obtain the beat type to which each first audio segment belongs; Based on the beat type to which each of the first audio segments belongs, the multiple audio transition points are determined.

13. The apparatus according to claim 11, characterized in that, The parsing module is used to segment the audio information to obtain multiple second audio segments; and to extract features from the multiple second audio segments to obtain the features of each second audio segment. Based on the characteristics of each second audio segment, the second audio segments are clustered to obtain at least one cluster, and each cluster includes at least one second audio segment; the plurality of audio inflection points are determined based on the at least one cluster.

14. The apparatus according to claim 9, characterized in that, The selection module is used to randomly generate a target camera movement number, select the target camera movement information corresponding to the target camera movement number from the plurality of candidate camera movement information, and each candidate camera movement information corresponds to a camera movement number; or, based on the operation of selecting any candidate camera movement information in the camera movement display list, the candidate camera movement information is determined as the target camera movement information, and the camera movement display list is used to display the plurality of candidate camera movement information.

15. A display device for a scene image, characterized in that, The device includes: The display module is used to display the live broadcast interface, which includes a first scene image. The first scene image is a picture obtained by a virtual camera capturing a virtual scene based on first camera movement information. The playback module is used to play background music; The switching module is used to switch the first camera movement information to the second camera movement information in response to the background music playing at the camera movement switching moment. The display module is also used to display a second scene on the live streaming interface. The second scene is a picture of the virtual scene captured by the virtual camera based on the second camera movement information. The camera movement switching time, the first camera movement information, and the second camera movement information are all determined by the apparatus according to any one of claims 9 to 14.

16. The apparatus according to claim 15, characterized in that, The switching module is used to display a camera movement switching control on the live streaming interface; and to switch the first camera movement information to the second camera movement information based on the confirmation operation of the camera movement switching control.

17. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by the processor to enable the electronic device to implement the method for determining camera movement information as described in any one of claims 1 to 6 or the method for displaying scene images as described in any one of claims 7 to 8.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to implement the method for determining camera movement information as described in any one of claims 1 to 6 or the method for displaying scene images as described in any one of claims 7 to 8.

19. A computer program product, characterized in that, The computer program product stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to implement the method for determining camera movement information as described in any one of claims 1 to 6 or the method for displaying scene images as described in any one of claims 7 to 8.

Citation Information

Patent Citations

  • Video shooting method and device, equipment and storage medium

    CN109922268A

  • System and method for generating word-by-word lyric file based on K-means clustering algorithm

    CN110867180A

  • Facial animation generation method and system and storage medium

    CN113609255A

  • Live broadcast control method and device, storage medium and electronic equipment

    CN114363689A