Information interaction processing method and device, computer equipment and storage medium
By displaying a virtual object that is displaced relative to the target object during video recording and responding to user voice control, the problem of poor interactive effects in traditional video recording is solved, enabling personalized and engaging video generation.
Patent Information
- Application Number
- CN202410594719.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional video recording cannot achieve real-time interaction, resulting in poor interactive effects.
During video recording, a virtual object with relative displacement to the target object is displayed, and the virtual object is controlled by voice to perform interactive actions. Recommended information containing prompts is displayed to encourage the user to input voice.
It improves the personalization of video content and user engagement, enhances users' interest in video recording, and enables the generation of unique video content.
Smart Images

Figure CN120935412A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an information interaction processing method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the development of technology, terminals can support more and more applications and have more and more powerful functions. People often use terminals in their daily life and production activities, such as using mobile phones to record videos.
[0003] Traditional video recording typically involves passively capturing the real world. Creators record the video footage through a terminal and interact with the content after recording. However, they cannot interact with the real-time video footage during the recording process, resulting in poor interactive effects on the video content. Summary of the Invention
[0004] Therefore, it is necessary to provide an information interaction processing method, apparatus, computer equipment, and storage medium that can improve the interaction effect with video content, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides an information interaction processing method. The method includes:
[0006] During video recording, a virtual object with a relative displacement to the target object is displayed;
[0007] Display recommended information containing prompts, which are used to prompt for voice input;
[0008] In response to receiving voice, the virtual object is displayed to perform an interactive action on the target object;
[0009] In response to the end of video recording, a video is obtained, the video containing the content displayed during the video recording process.
[0010] Secondly, this application also provides an information interaction processing apparatus. The apparatus includes:
[0011] The virtual object display module is used to display virtual objects that have a relative displacement to the target object during video recording.
[0012] The recommendation information display module is used to display recommendation information containing prompts, the prompts being used to prompt for voice input;
[0013] The voice control module is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object;
[0014] The video acquisition module is used to acquire a video in response to the end of video recording, the video containing the content displayed on the video recording screen during the video recording process.
[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the methods described above.
[0016] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of any of the methods described above.
[0017] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of any of the methods described above.
[0018] The aforementioned information interaction processing method, apparatus, computer equipment, storage medium, and computer program product, during video recording, display an interactive virtual object with relative displacement to the target object. Users can observe the positional changes of the virtual object relative to the target object. By displaying recommended information containing prompts, users can be prompted to input voice to control the interaction between the virtual object and the target object. In response to receiving voice input, the virtual object executes interactive actions against the target object. This allows users to interact with the virtual object and target object via voice control during video recording. At the end of recording, a video recording the interaction between the virtual object and the target object based on voice is obtained. By allowing users to control the interaction between the virtual object and the target object via voice during video recording, videos generated by different users are unique, and the content of each video recorded by the same user is also different, enhancing user participation and control over the video content and further improving the personalization of the video content. Simultaneously, the interaction results of the virtual object against the target object demonstrate the accuracy of the user's voice control of the virtual object's interactive actions, thereby increasing user interest in participating in video recording.
[0019] In addition, an information interaction processing method, apparatus, computer equipment, and storage medium are also provided.
[0020] Sixthly, this application provides an information interaction processing method. The method includes:
[0021] Displays a virtual object that has a relative displacement to the target object;
[0022] Display recommended information containing prompts, which are used to prompt for voice input;
[0023] In response to receiving voice, the virtual object is displayed to perform an interactive action on the target object;
[0024] In response to the triggering operation of the recommended information, the recommended information interface is displayed.
[0025] Seventhly, this application also provides an information interaction processing apparatus. The apparatus includes:
[0026] The virtual object display module is used to display virtual objects that have a relative displacement with respect to the target object;
[0027] The recommendation information display module is used to display recommendation information containing prompts, the prompts being used to prompt for voice input;
[0028] The voice control module is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object;
[0029] The recommendation interface display module is used to display the recommendation information interface in response to a trigger operation on the recommendation information.
[0030] Eighthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the methods described above.
[0031] Ninthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of any of the methods described below.
[0032] In a tenth aspect, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of any of the methods described above.
[0033] The aforementioned information interaction processing method, apparatus, computer equipment, storage medium, and computer program product, by displaying an interactive virtual object with relative displacement to the target object, allow the user to see the positional changes of the virtual object relative to the target object. By displaying recommended information containing prompts, the user can be prompted to input voice to control the interaction between the virtual object and the target object. In response to receiving voice input, the virtual object executes an interactive action against the target object. In response to triggering the recommended information, a recommended information interface is displayed. This enables users to interact with the virtual object and target object via voice control, and simultaneously interact with the recommended information to obtain more information. It avoids the situation where users can only switch to other interfaces to obtain more information related to the recommended information after the interaction between the virtual object and the target object is complete, thus improving information acquisition efficiency. Attached Figure Description
[0034] Figure 1 This is an application environment diagram of an information interaction processing method in one embodiment;
[0035] Figure 2 This is a flowchart illustrating an information interaction processing method in one embodiment;
[0036] Figure 3 This is a schematic diagram of a video recording screen in one embodiment;
[0037] Figure 4 This is a schematic diagram of a video recording screen in another embodiment;
[0038] Figure 5 This is a schematic diagram of a video recording screen in another embodiment;
[0039] Figure 6 This is a schematic diagram of a video recording page in one embodiment;
[0040] Figure 7 This is a schematic diagram of a video recording page in another embodiment;
[0041] Figure 8 This is a schematic diagram of a video recording page in another embodiment;
[0042] Figure 9 This is a schematic diagram illustrating the effect of the avoidance action in one embodiment;
[0043] Figure 10 This is a schematic diagram illustrating the effect of the avoidance action in another embodiment;
[0044] Figure 11 This is a schematic diagram of the avoidance area corresponding to a virtual object in one embodiment;
[0045] Figure 12This is a schematic diagram of a virtual object in one embodiment;
[0046] Figure 13 This is a schematic diagram of a synthesized virtual object in one embodiment;
[0047] Figure 14 This is a flowchart illustrating the information interaction processing method in another embodiment.
[0048] Figure 15 This is a flowchart illustrating the information interaction processing method in another embodiment;
[0049] Figure 16 This is a structural block diagram of an information interaction processing device in one embodiment;
[0050] Figure 17 This is a structural block diagram of the information interaction processing device in another embodiment;
[0051] Figure 18 This is a structural block diagram of the information interaction processing device in another embodiment;
[0052] Figure 19 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0054] The information interaction processing method provided in this application involves artificial intelligence technologies such as machine learning and computer vision, wherein:
[0055] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0056] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0057] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the vision field, such as Swin-transformer, ViT, V-MOE, and MAE, can be quickly and widely applied to specific downstream tasks after fine-tuning. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and other technologies. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0058] The information interaction processing method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be placed in the cloud or on another server. This information interaction processing method can be executed independently by terminal 102, or it can be executed interactively with server 104. Taking terminal 102 executing independently as an example, during video recording, terminal 102 displays a virtual object with a relative displacement to the target object; displays recommended information containing prompts to encourage voice input; in response to receiving voice input, displays the virtual object to perform interactive actions targeting the target object; and in response to the end of video recording, obtains the video, which includes the content displayed during the recording process.
[0059] The terminals can be, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, portable wearable devices, and network devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Network devices can include routers, switches, firewalls, load balancers, network storage devices, network adapters, etc.
[0060] Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication; this application does not impose any restrictions on this.
[0061] In one embodiment, the information interaction processing method is performed by a computer device, which may be, for example, a computer device that can be... Figure 1 The terminal 102 or server 104 shown is used by this interaction method. Figure 1 The following explanation uses the terminal 102 shown as an example. Figure 2 As shown, the method may include the following steps:
[0062] S202, During video recording, a virtual object with relative displacement to the target object is displayed.
[0063] Video recording refers to the process of capturing and saving real-time video and audio through a terminal. This real-time video is the video recording screen. In this embodiment, the video recording screen is an interactive screen displayed by the terminal, meaning the user can interact with it in real time. Specifically, this interaction can be with virtual objects displayed on the video recording screen; for example, controlling the behavior of virtual objects through interaction. The video recording screen can be footage captured from a real environment, or it can be footage generated or synthesized through various methods, such as augmented reality, virtual reality, or mixed reality.
[0064] Virtual objects are graphics, models, or entities created in a digital environment. They can be generated through computer graphics technology and can take various forms, such as virtual characters or virtual animals.
[0065] The target object is an element used to obstruct the movement of a virtual object. It can be generated through computer graphics technology and can take various forms, such as virtual vehicles, virtual railings, virtual animals, virtual characters, virtual trees, virtual ditches, etc.
[0066] In one embodiment, in response to the start operation of video recording, the terminal displays a video recording screen and shows a virtual object that has a relative displacement with the target object according to a preset initial moving speed. After the virtual object successfully interacts with the target object, the video recording screen displays the movement state of the virtual object with an adjusted moving speed relative to other target objects, wherein the adjusted moving speed may be the same as or different from the initial moving speed.
[0067] like Figure 3 The image shown is a schematic diagram of a video recording screen in one embodiment. The video recording screen displays a virtual object 302 and a target object (railing) 304. In this video recording screen, the virtual object 302 moves towards the target object (railing) 304 at an initial moving speed. After the virtual object successfully avoids the target object (railing) 304 by performing an avoidance maneuver, ... Figure 4 The video recording shown depicts the virtual object 302 moving towards the target object (railing) 306 at an adjusted speed.
[0068] S204, Display recommended information including prompts.
[0069] The prompts are used to guide voice input, while the recommendations are used to recommend specific products, services, or brands. For example... Figure 3The recommended message shown in the video recording is "XX Milk is so delicious, shout 'XX Milk' to complete the hurdles." Here, XX Milk is a milk brand, and the recommended message is for XX Milk. "Shout 'XX Milk' to complete the hurdles" is the prompt message included in the recommended message.
[0070] S206, In response to receiving the voice, display the virtual object to perform an interactive action on the target object.
[0071] Speech refers to sound signals produced through vocal organs (such as the mouth and vocal cords) for communication and information transmission; specifically, it can be speaking or singing.
[0072] Interactive actions are the responses of virtual objects to voice commands input by users. Specifically, interactive actions can be movement actions targeting a target object, such as moving to a specific location, or avoidance actions. Avoidance actions aim to avoid target objects that hinder the movement of virtual objects, such as jumping upwards, moving to the left, or moving to the right.
[0073] In one embodiment, during video recording, the terminal can record sound in the environment of the video recording using an audio recording device to obtain recorded audio. Real-time speech recognition is then performed on the recorded audio. When speech is detected in the video recording environment, speech is extracted from the recorded audio to obtain speech segments. Information recognition is then performed on these speech segments to obtain speech information. It is then determined whether the obtained speech information matches the pre-configured interaction trigger information for the target object. If it matches, a virtual object is displayed on the video recording screen to perform an interactive action on the target object. If it does not match, the virtual object is displayed on the video recording screen maintaining a relative displacement with the target object until a collision occurs. Alternatively, when speech is detected again in the video recording environment, the speech information is extracted, and the step of determining whether the obtained speech information matches the pre-configured interaction trigger information for the target object is repeated.
[0074] Speech information refers to specific aspects of speech characteristics, which can be at least one of content information and speech attribute information. Content information, also known as semantic information, refers to the actual meaning or connotation conveyed in speech, i.e., text content. Speech attribute information refers to the characteristics and properties of the sound itself, which can be at least one of speech volume, speech pitch, and speech rhythm. These attributes reflect the characteristics of the speaker's sound wave amplitude, frequency changes, and rhythmic patterns during the process of vocalization.
[0075] Interactive trigger information is pre-configured information used to trigger virtual objects to perform specific actions. Specifically, it can be at least one of the following sub-voice trigger information: volume trigger information, tone trigger information, and content trigger information. Specifically, volume trigger information can be a volume trigger threshold, tone trigger information can be a preset voice tone, and content trigger information can be preset voice content. Interactive trigger information can also be song trigger information from song rhythm trigger information, song melody trigger information, and song content trigger information. Specifically, song rhythm trigger information can be the rhythm information of the corresponding song segment of the target song, song melody trigger information can be the melody information of the corresponding song segment of the target song, and song content trigger information can be the lyrics information of the corresponding song segment of the target song.
[0076] Matching conditions are used to determine whether the actual voice input received in the recording video environment matches the preset interaction trigger information. For example, the content trigger information might be keywords like "left," "right," or "jump." In a real-world scenario, when the similarity between the user's spoken words and a preset keyword reaches a content similarity trigger threshold, the voice information is determined to match the pre-configured interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. Conversely, if the similarity between the user's spoken words and any preset keyword does not reach the similarity trigger threshold, the voice information is determined to not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action. Another example is the volume trigger information, which is a certain volume threshold. In a real-world scenario, when the user's voice volume reaches this threshold, the voice input is determined to match the preset interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. If the information matches the pre-configured interaction trigger information for the target object, the virtual object can be triggered to perform the corresponding action. If the volume of the user's voice does not reach the volume threshold, it is determined that the voice information does not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action. For example, if the tone trigger information is a preset voice tone, if the similarity between the user's voice tone and the preset voice tone reaches the tone similarity trigger threshold, it is determined that the voice information matches the pre-configured interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. If the similarity between the user's voice tone and the preset voice tone does not reach the tone similarity trigger threshold, it is determined that the voice information does not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action.
[0077] refer to Figure 3The video recording screen displays a virtual object 302 and a target object (railing) 304. In this video recording screen, the virtual object 302 moves towards the target object (railing) 304 at an initial moving speed. During the video recording process before the virtual object 302 moves to the target object (railing) 304, when voice is emitted in the recording environment, the voice information responding to the voice matches the pre-configured interaction trigger information for the target object. Figure 5 The video recording shown depicts a virtual object 302 performing an avoidance maneuver (hurdle jump) against a target object (railing) 304. After the virtual object 302 successfully avoids the target object (railing) 304, Figure 4 The video recording shown depicts the virtual object 302 moving toward the target object (railing) 306 at an adjusted speed. During the video recording process before the virtual object 302 moves to the target object (railing) 306, step S206 is re-executed.
[0078] S208, in response to the end of video recording, obtain the video, which includes the content displayed on the video recording screen during the video recording process.
[0079] The end of video recording can be triggered in different ways, including command triggering or event triggering. Command triggering can be triggered by the terminal detecting a click on the recording stop control or detecting a user's voice command to stop recording. Event triggering can be triggered by automatically stopping when a specific event occurs, such as reaching a preset time length, the virtual object failing to avoid the target object, or the virtual object successfully avoiding all target objects.
[0080] Specifically, when the terminal detects a stop command or a specific event, it terminates the currently ongoing video recording process, such as stopping the capture of video recording frames, encoding the content displayed in the captured video recording frames to convert it into a specific file format, such as MP4 or AVI, and saving the video file containing the content displayed in the video recording frames throughout the entire recording process.
[0081] In one embodiment, the terminal can also preview and play the acquired video so that the user can check or edit the acquired video.
[0082] refer to Figure 6 The video recording page shown displays a recording control 602 and a preview screen 604. The user can click the recording control 602, and the terminal will begin recording video in response to the click. Figure 7As shown, a video recording screen 702 is simultaneously displayed on the video recording page. The video recording screen 702 displays a virtual object moving towards the target object and operation guidance information 704. The user can issue a voice command based on the operation guidance information 704 as the virtual object moves towards the target object. If the terminal's voice response matches the pre-configured interaction trigger information for the target object, the video recording screen displays the virtual object performing an avoidance action towards the target object. Figure 8 In the hurdle jumping action, after the virtual object successfully avoids the target object (hurdle) by performing an avoidance action, the virtual object is displayed in the video recording screen moving towards the next target object (hurdle) at an adjusted speed. During the video recording process before the virtual object moves to the next target object (hurdle), step S206 is repeated. When the user wants to stop the video recording, they can click... Figure 8 The recording control 802 in the terminal responds to the trigger operation of the recording control 802 to end the video recording and obtain the video containing the content displayed on the video recording screen during the video recording process.
[0083] In the aforementioned information interaction processing method, during video recording, an interactive virtual object with relative displacement to the target object is displayed. Users can observe the positional changes of the virtual object relative to the target object. Recommended information containing prompts can encourage users to input voice commands to control the interaction between the virtual object and the target object. In response to voice input, the virtual object executes interactive actions against the target object. This allows users to interact with the virtual object and target object via voice control during video recording. At the end of recording, a video recording the interaction between the virtual object and the target object based on voice is obtained. By allowing users to control the interaction between the virtual object and the target object via voice during video recording, videos generated by different users are unique, and the content of each video recorded by the same user is also different, enhancing user participation and control over the video content and further improving the personalization of the video content. Simultaneously, the interaction results of the virtual object against the target object demonstrate the accuracy of the user's voice control of the virtual object's interactive actions, thereby increasing user interest in participating in video recording.
[0084] In one embodiment, the process of a terminal displaying a virtual object performing an interactive action against a target object includes the following steps: on the video recording screen, the virtual object is displayed performing an avoidance action against the target object according to the action amplitude associated with the voice information; wherein, when the action amplitude meets a preset amplitude condition, the virtual object successfully avoids the target object.
[0085] Among them, the motion amplitude refers to the degree or range exhibited by the virtual object when performing an avoidance action. For example, when the avoidance action is an upward hurdle jump, the corresponding motion amplitude can be the height of the jump or the vertical distance. For example, the height of the virtual object's jump can be 100 pixels. When the avoidance action is a horizontal movement to the left or right, the corresponding motion amplitude can be the distance of the horizontal movement to the left or right. For example, the distance of the virtual object moving to the left can be 50. When the avoidance action is a rotational action around its own axis or a certain point, the corresponding motion amplitude can be the angle or direction of rotation, which can be expressed in degrees. For example, the virtual object rotates 90 degrees to the left around its own axis.
[0086] It can be understood that the motion amplitude is the ideal degree or range of motion of a virtual object when performing a specific action without being affected by a target object. In other words, it is the range or degree of motion that a virtual object should ideally achieve. For example, if a virtual object performs an upward jump, the motion amplitude can represent the ideal jump height of the virtual object, that is, the maximum height that the virtual object should jump when there is no target object. When the virtual object performs an avoidance action according to this motion amplitude, if it comes into contact with a target object, that is, is affected by the target object, the execution process of the avoidance action will change. For example, in the event of a severe collision, the execution process of the avoidance action may be interrupted, thus making it impossible to complete the avoidance action according to the motion amplitude. Or, in the event of a slight contact, the trajectory of the virtual object will change, thus making it impossible to reach the expected jump height.
[0087] The specific range of motion associated with voice information can be adjusted based on the voice information itself. Voice information can include at least one sub-voice information such as voice volume, voice tone, or voice content. Voice information can also include at least one song information such as the rhythm, melody, or content of a sung song. For example, when the voice information is voice volume, the corresponding motion range can be adjusted based on the volume level. When the voice information is voice content, the corresponding motion range can be adjusted based on the similarity between the voice content and the content trigger information. When the voice information is voice tone, the corresponding motion range can be adjusted based on the similarity between the voice tone and a preset voice tone. When the voice information is the rhythm of a sung song, the corresponding motion range can be adjusted based on the similarity between the rhythm of the sung song and the standard rhythm of the target song. When the voice information is the melody of a sung song, the corresponding motion range can be adjusted based on the similarity between the melody of the sung song and the standard melody of the target song. When the voice information is the content of a sung song, the corresponding motion range can be adjusted based on the similarity between the content of the sung song and the standard content of the target song. Here, "sung song" refers to the version of the song expressed by the user through singing or chanting the target song, and "target song" refers to the original recorded version of the song, also known as the original version.
[0088] The preset amplitude condition is a pre-defined range of motion that a virtual object must achieve to successfully avoid a target object. In other words, if the range of motion of the virtual object when performing the avoidance action meets the preset amplitude condition, then the virtual object can successfully avoid the target object, that is, the avoidance action is successful. Conversely, if the range of motion does not meet the preset amplitude condition, that is, it does not reach the required level or amplitude, then the virtual object's avoidance action fails, which may result in a collision or occlusion with the target object.
[0089] The motion amplitude condition can be specifically a motion amplitude threshold condition. A threshold or range can be set. When the motion amplitude of the virtual object reaches the threshold or is within the range, the motion amplitude of the virtual object is determined to meet the preset amplitude condition.
[0090] Specifically, after receiving voice information, the terminal can determine the corresponding action amplitude of the virtual object based on the voice information and the amplitude function corresponding to the voice information. The virtual object is then displayed on the video recording screen to perform an avoidance action against the target object according to the determined action amplitude. For example, the motion trajectory and posture of the virtual object can be determined based on the determined action amplitude, the distance between the virtual object and the target object, and the movement speed of the virtual object. The motion trajectory and posture of the virtual object are displayed on the video recording screen to show the execution effect of the avoidance action. When the action amplitude of the virtual object meets the preset amplitude conditions, the execution effect of the avoidance action of the virtual object successfully avoiding the target object is shown. When the action amplitude of the virtual object does not meet the preset amplitude conditions, the execution effect of the avoidance action of the virtual object colliding with the target object is shown.
[0091] In the above embodiments, the terminal controls the virtual object to perform avoidance actions against the target object and the range of motion based on the avoidance actions through voice information, so that the user can participate more intuitively in the video recording process, improve the user's participation and control over the video content, and thus further improve the personalization of the video content.
[0092] In one embodiment, the preset amplitude conditions include a first amplitude condition and a second amplitude condition. When the amplitude of the action meets the preset amplitude conditions, the process in which the virtual object successfully avoids the target object includes the following steps: when the amplitude of the action meets the first amplitude condition, the virtual object successfully avoids the target object in a non-contact manner; when the amplitude of the action meets the second amplitude condition, the virtual object successfully avoids the target object in a contact manner.
[0093] The first amplitude condition is a pre-defined amplitude condition that a virtual object must meet to successfully avoid a target object without contact. The second amplitude condition is a pre-defined amplitude condition that a virtual object must meet to successfully avoid a target object with contact. It can be understood that if the amplitude of the virtual object's avoidance action meets the first amplitude condition, then the virtual object can successfully avoid the target object without contact. If the amplitude of the virtual object's avoidance action meets the second amplitude condition, then the virtual object can successfully avoid the target object with slight contact.
[0094] The first motion amplitude condition can be specifically set to a first amplitude range, and the second motion amplitude condition can be specifically set to a second amplitude range. The first amplitude range and the second amplitude range do not overlap, and the boundary value of the second amplitude range is less than the boundary value of the first amplitude range. When the motion amplitude of the virtual object is within the first amplitude range, it is determined that the motion amplitude of the virtual object satisfies the first motion amplitude condition. When the motion amplitude of the virtual object is within the second amplitude range, it is determined that the motion amplitude of the virtual object satisfies the second motion amplitude condition.
[0095] Specifically, after determining the range of motion of the virtual object, the terminal displays on the video recording screen that the virtual object performs an avoidance action against the target object according to the determined range of motion. When the range of motion of the virtual object meets the first range condition, the execution effect of the virtual object successfully avoiding the target object without contact is displayed. When the range of motion of the virtual object meets the second range condition, the execution effect of the virtual object successfully avoiding the target object with slight contact is displayed. When the range of motion of the virtual object does not meet the first range condition and does not meet the second range condition, the execution effect of the virtual object colliding with the target object is displayed, that is, the virtual object fails to avoid the target object.
[0096] by Figure 9 The three scenarios shown illustrate the above embodiments. After determining the amplitude of the virtual object's movement based on the voice information, the terminal can further determine the amplitude range of that movement. When the value of the movement amplitude is greater than b and less than or equal to a, it is determined that the movement amplitude is within amplitude range 1, i.e., the first amplitude condition is met, and then it is displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object in a non-contact manner, as shown in scenario 2; when the value of the action amplitude is greater than c and less than or equal to b, the action amplitude is determined to be within amplitude range 2, that is, the second amplitude condition is met, and then it is displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 3; when the value of the action amplitude is greater than d and less than or equal to c, it is determined that the action amplitude is within amplitude range 3, that is, it does not meet the first amplitude condition and does not meet the second amplitude condition, and is then displayed on the video recording screen. Figure 9 The execution effect of the virtual object colliding with the target object as shown in scenario 1.
[0097] In the above embodiments, the terminal associates the range of motion with the avoidance result achieved by the virtual object, thereby enabling the user to control the range of motion of the virtual object to perform avoidance actions through voice information, thereby achieving different avoidance results. This allows the user to more intuitively understand the avoidance effect corresponding to different ranges of motion and realizes personalized avoidance strategies, while enhancing the user's participation and control in the interaction process.
[0098] In one embodiment, the voice information includes at least one sub-voice information among voice volume, voice tone, or voice content, and the action amplitude includes a first action amplitude corresponding to at least one sub-voice information; the process of the terminal displaying the virtual object performing an avoidance action against the target object according to the action amplitude associated with the voice information on the video recording screen includes the following steps: displaying the virtual object performing an avoidance action against the target object according to the first action amplitude on the video recording screen.
[0099] Understandably, when the first action amplitude corresponds to only one type of sub-speech information, it can be directly determined based solely on the sub-amplitude corresponding to that associated sub-speech information. When the first action amplitude corresponds to two or more types of sub-speech information, it can be determined by weighted summation of the sub-amplitudes corresponding to each associated sub-speech information. For example, if the first action amplitude corresponds to both speech volume and speech content, the volume sub-amplitude can be determined based on the speech volume, and the content sub-amplitude can be determined based on the speech content. The volume weight corresponding to the speech volume and the content weight corresponding to the speech content can be obtained, and the volume sub-amplitude and content sub-amplitude can be weighted and summed based on the volume weight and content weight to obtain the first action amplitude. If the speech volume is more important than the speech content, then a larger weight value can be given to the speech volume. In this way, the sub-amplitude corresponding to the speech volume will have a larger proportion when determining the first action amplitude, thus having a greater impact on the final first action amplitude.
[0100] Specifically, after obtaining the target sub-speech information in the speech volume, tone, or content associated with the first action amplitude, the terminal determines the first action amplitude based on the target sub-speech information and the amplitude function corresponding to the target sub-speech information. The virtual object is then displayed on the video recording screen to perform an avoidance action against the target object according to the determined first action amplitude. For example, the motion trajectory and posture of the virtual object can be determined based on the determined first action amplitude, the distance between the virtual object and the target object, and the movement speed of the virtual object. The motion trajectory and posture of the virtual object are then displayed on the video recording screen to present the execution effect of the avoidance action. When the first action amplitude of the virtual object meets the preset amplitude condition, the execution effect of the avoidance action of the virtual object successfully avoiding the target object is presented. When the first action amplitude of the virtual object does not meet the preset amplitude condition, the execution effect of the avoidance action of the virtual object colliding with the target object is presented.
[0101] Taking voice volume as an example, the above embodiment is explained as follows: After the terminal extracts the voice volume, it compares the voice volume with a volume trigger threshold. Assuming the volume trigger threshold is 50 decibels, when the voice volume is less than or equal to 50 decibels, the virtual object will not be triggered to perform an avoidance action; when the voice volume is greater than 50 decibels, a corresponding first action amplitude is determined based on the voice volume. The higher the voice volume, the larger the corresponding action amplitude. When the voice volume is greater than 50 decibels and less than or equal to 60 decibels, the determined first action amplitude corresponds to... Figure 9 If the amplitude range 3 in the video recording is not satisfied, meaning neither the first amplitude condition nor the second amplitude condition is met, then the video recording screen will display... Figure 9 The execution effect of the virtual object collision avoidance action shown in scenario 1; when the voice volume is greater than 60 decibels and less than or equal to 75 decibels, the determined first action amplitude corresponds to Figure 9 If the amplitude range 2 in the video recording meets the second amplitude condition, then it will be displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 3; when the voice volume is greater than 75 decibels, the determined first action amplitude corresponds to Figure 9 If the amplitude range is 1, that is, if the first amplitude condition is met, then it will be displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 2.
[0102] In the above embodiments, the terminal controls the range of the avoidance action by using at least one sub-voice information, such as voice volume, voice tone, or voice content. Users do not need complicated gestures or operations, which improves the convenience of users participating in video content. At the same time, it enriches the diversity of user interaction forms with video content, enhances the interactive fun between users and video content, and improves users' participation and control over video content, thereby further improving the personalization of video content.
[0103] In one embodiment, the voice is a song sung by the user according to the target song, and the voice information includes at least one of the song information such as the song rhythm, song melody, or song content. The action range includes a second action range corresponding to at least one song information. The process of the terminal displaying the virtual object performing an avoidance action against the target object according to the action range associated with the voice information on the video recording screen includes the following steps: displaying the virtual object performing an avoidance action against the target object according to the second action range on the video recording screen.
[0104] Among them, the user can be the user recording the video, and the target song can be the song that the user is expected to sing during the video recording process. The target song can be a system default song or a song selected by the user on the video recording page before the video recording begins.
[0105] For example, suppose a user wants to record a video, and the background music is a popular song that the system defaults to. In this case, the target song is the popular song that the user will sing during the video recording. During the recording process, the user can control the virtual object by singing the target song. Alternatively, if the user wants to sing their favorite song during the video recording, they can select their favorite song as the target song on the video recording page. In this way, the user can control the virtual object by singing the target song during the video recording process.
[0106] Understandably, when the second action amplitude corresponds to only one type of song information, it can be directly determined based solely on the sub-amplitude corresponding to the associated song information. When the second action amplitude corresponds to two or more types of song information, it can be determined by weighted summation of the sub-amplitudes corresponding to each associated song information. For example, if the second action amplitude corresponds to both song rhythm and song melody, the rhythm sub-amplitude can be determined based on the song rhythm, and the melody sub-amplitude based on the song melody. The rhythm weight and melody weight corresponding to the song rhythm and melody can then be obtained, and the rhythm and melody sub-amplitudes can be weighted and summed to obtain the second action amplitude. If the song rhythm is more important than the song melody, then a larger weight can be given to the song rhythm. In this way, the sub-amplitude corresponding to the song rhythm will have a larger proportion when determining the second action amplitude, thus having a greater impact on the final second action amplitude.
[0107] Specifically, after obtaining the target song information from the song rhythm, voice pitch, or melody associated with the second action amplitude, the terminal determines the second action amplitude based on the target song information and the amplitude function corresponding to the target song information. The virtual object is then displayed on the video recording screen performing an avoidance action against the target object according to the determined second action amplitude. For example, the trajectory and posture of the virtual object can be determined based on the determined second action amplitude, the distance between the virtual object and the target object, and the movement speed of the virtual object. The trajectory and posture of the virtual object are then displayed on the video recording screen to present the effect of the avoidance action. When the second action amplitude of the virtual object meets the preset amplitude conditions, the effect of the virtual object successfully avoiding the target object is presented. When the second action amplitude of the virtual object does not meet the preset amplitude conditions, the effect of the virtual object colliding with the target object is presented.
[0108] Taking the melody of a sung song as an example, the above embodiment is explained as follows: After the terminal extracts the melody of the sung song, it determines the similarity score between the melody of the sung song and the standard melody of the target song. Assuming the trigger similarity threshold is 50 points, when the similarity score is less than or equal to 50 points, the virtual object will not be triggered to perform an avoidance action; when the similarity score is greater than 50 points, the corresponding first action amplitude is determined based on the similarity score. The higher the similarity score, the larger the corresponding action amplitude. When the similarity score is greater than 50 points and less than or equal to 60 points, the determined first action amplitude corresponds to... Figure 9 If the amplitude range 3 in the video recording is not satisfied, meaning neither the first amplitude condition nor the second amplitude condition is met, then the video recording screen will display... Figure 9The execution effect of the virtual object collision avoidance action shown in scenario 1; when the similarity score is greater than 60 points and less than or equal to 80 points, the determined first action amplitude corresponds to Figure 9 If the amplitude range 2 in the video recording meets the second amplitude condition, then it will be displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 3; when the similarity score is greater than 80 and less than or equal to 100, the determined first action amplitude corresponds to Figure 9 If the amplitude range is 1, that is, if the first amplitude condition is met, then it will be displayed on the video recording screen. Figure 9 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 2.
[0109] In the above embodiments, the terminal sings a song, thereby controlling the range of the avoidance action based on at least one piece of song information, such as the song's rhythm, melody, or content. This enriches the diversity of user interaction with video content, enhances the interactive fun between users and video content, and improves users' participation and control over video content, thereby further enhancing the personalization of video content.
[0110] In one embodiment, the above information interaction processing method further includes the following steps: playing the accompaniment of the target song during video recording; and, in response to obtaining the voice, displaying the similarity between at least one song information and the standard song information of the target song on the video recording screen.
[0111] Among them, the accompaniment to a song refers to the audio track of a song played without a singer, which usually includes instrumental music and rhythm, but does not include the lyrics.
[0112] Specifically, the terminal starts playing the accompaniment to the song while starting video recording to ensure that the audio playback and video recording are synchronized and to avoid audio-visual asynchrony. At the same time, it collects the sound emitted in the environment of the video recording through the microphone. When speech is detected in the environment of the video recording, the recorded speech is recognized to obtain at least one of the song information, such as the song rhythm, song melody, or song content. After obtaining the target song information in the song rhythm, speech pitch, or song melody associated with the second action amplitude, the similarity between the target song information and the standard song information corresponding to the target song is determined, and the similarity information is dynamically displayed on the video recording screen, for example, through a progress bar, percentage numbers, or graphical methods.
[0113] like Figure 10As shown, taking the melody of a sung song as an example, the above embodiment is explained. After the terminal extracts the melody of the sung song, it determines the similarity score between the melody of the sung song and the standard melody of the target song. Assuming the trigger similarity threshold is 50 points, when the similarity score is less than or equal to 50 points, the virtual object will not be triggered to perform an avoidance action; when the similarity score is greater than 50 points, the corresponding first action amplitude is determined according to the similarity score. The higher the similarity score, the larger the corresponding action amplitude. When the similarity score is greater than 50 points and less than or equal to 60 points, the determined first action amplitude corresponds to... Figure 10 If the amplitude range 3 in the video recording is not satisfied, meaning neither the first amplitude condition nor the second amplitude condition is met, then the video recording screen will display... Figure 10 The execution effect of the virtual object collision avoidance action shown in scenario 1; when the similarity score is greater than 60 points and less than or equal to 80 points, the determined first action amplitude corresponds to Figure 10 If the amplitude range 2 in the video recording meets the second amplitude condition, then it will be displayed on the video recording screen. Figure 10 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 3; when the similarity score is greater than 80 and less than or equal to 100, the determined first action amplitude corresponds to Figure 10 If the amplitude range is 1, that is, if the first amplitude condition is met, then it will be displayed on the video recording screen. Figure 10 The execution effect of the virtual object successfully avoiding the target object through contact, as shown in scenario 2.
[0114] In the above embodiments, by playing the accompaniment of the target song during video recording, the terminal can help users better understand the rhythm, melody, and emotional expression of the song, thereby improving the accuracy and quality of their singing. This also helps users increase the similarity between the song information of their singing and the standard song information of the target song, thus enabling more accurate control of virtual objects to perform avoidance actions. Furthermore, by displaying the similarity between the user's singing and the standard song information of the target song on the video recording screen, users can better understand the degree of closeness between their singing and the target song, allowing them to adjust their singing style to improve the similarity between the singing and the standard song information, thereby enhancing the control effect on virtual objects to perform avoidance actions.
[0115] In one embodiment, the above information interaction processing method further includes the following steps: when the target object is successfully avoided in a contactless manner, the accompaniment of the next song segment of the target song is played; wherein, the playback speed of the accompaniment of the next song segment is greater than the original playback speed, and the original playback speed refers to the playback speed of the accompaniment of the song during the video recording process before the virtual object avoids the current target object.
[0116] Understandably, after confirming that the virtual object successfully avoids the target object without contact, the difficulty of subsequent virtual objects avoiding the target object can be increased. Specifically, the playback speed of the song accompaniment can be increased. By speeding up the rhythm and beat of the music, the user's reaction speed and control over the virtual object can be enhanced, thereby increasing the challenge of avoiding the target object. The increased difficulty will stimulate the user's desire to challenge, thereby increasing the user's interest in participating in video recording.
[0117] In one embodiment, the process of a terminal displaying a virtual object performing an avoidance action against a target object on a video recording screen includes the following steps: on the video recording screen, displaying a virtual object performing an avoidance action against a target object according to a fixed movement amplitude; wherein, if the voice occurs when the virtual object is in the avoidable area of the target object in the video screen, the virtual object successfully avoids the target object.
[0118] Among them, the fixed motion amplitude is a pre-set, fixed and unchanging degree or range exhibited by the virtual object when performing an avoidance action, and its size is not affected by voice information.
[0119] The avoidable area of the target object refers to the area in which the virtual object can successfully avoid the target object when performing an avoidance action with a fixed action range. It can be preset according to the shape, size, fixed action range, etc. of the target object.
[0120] Specifically, when voice is emitted in the video recording environment, the terminal responds to the voice information and determines whether the voice occurs when the virtual object is in the avoidable area of the target object in the video frame, if the voice information matches the pre-configured interaction trigger information for the target object. If so, the terminal obtains the preset fixed motion amplitude and displays the virtual object performing the avoidance action against the target object according to the fixed motion amplitude on the video recording screen, presenting the effect of the virtual object successfully avoiding the target object.
[0121] In the above embodiments, the terminal sets a dodgeable area corresponding to the target object, allowing the user to issue voice commands at appropriate times based on the relative positional relationship between the virtual object and the target object. This guides the virtual object to perform dodge actions. By issuing voice commands in a timely and accurate manner, the user can effectively control the behavior of the virtual object to successfully avoid the target object. This interaction method not only enhances the user's participation and control capabilities but also improves the real-time performance and accuracy of the interaction. Furthermore, the successful avoidance of the target object by the virtual object demonstrates the accuracy of the user's control over the timing of issuing voice commands, thereby increasing the user's interest in participating in video recording.
[0122] In one embodiment, the avoidable area includes a first avoidable area and a second avoidable area. If the voice occurs when the virtual object is within the avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in the following ways: if the voice occurs when the virtual object is within the first avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in a non-contact manner; if the voice occurs when the virtual object is within the second avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in a contact manner.
[0123] The first avoidable area can also be called the non-contact avoidable area, and the second avoidable area can also be called the contact avoidable area. The first avoidable area refers to the area in which the virtual object can successfully avoid the target object without contact when performing an avoidable action according to a fixed action range. The second avoidable area refers to the area in which the virtual object can successfully avoid the target object with slight contact when performing an avoidable action according to a fixed action range.
[0124] Specifically, when voice is emitted in the video recording environment, the terminal responds by determining whether the voice information matches the pre-configured interaction trigger information for the target object. If so, it determines whether the voice occurred when the virtual object was in the first avoidable area of the target object in the video frame. If so, it obtains a preset fixed motion amplitude and displays the virtual object performing an avoidance action against the target object according to the fixed motion amplitude on the video recording screen, presenting the effect of the virtual object successfully avoiding the target object without contact. If not, it determines whether the voice occurred when the virtual object was in the second avoidable area of the target object in the video frame. If so, it obtains a preset fixed motion amplitude and displays the virtual object performing an avoidance action against the target object according to the fixed motion amplitude on the video recording screen, presenting the effect of the virtual object successfully avoiding the target object with slight contact.
[0125] In the above embodiments, the terminal sets a first avoidable area and a second avoidable area corresponding to the target object, so that the user can choose to issue a voice command at an appropriate time according to the relative positional relationship between the virtual object and the target object, thereby guiding the virtual object to perform an avoidance action. By issuing voice commands in a timely and accurate manner, the user can effectively control the behavior of the virtual object and make it successfully avoid the target object. This interaction method not only enhances the user's participation and control capabilities, but also improves the real-time performance and accuracy of the interaction. At the same time, it can also demonstrate the accuracy of the user's control over the timing of issuing voice commands by showing how the virtual object successfully avoids the target object, thereby increasing the user's interest in participating in video recording.
[0126] In one embodiment, the above information interaction processing method further includes the following steps: if the voice occurs when the virtual object is in an unavoidable area of the target object in the video frame, the virtual object collides with the target object; if the voice occurs when the virtual object is in an invalid avoidance area of the target object in the video frame, after performing the avoidance action, the virtual object maintains a movement state with relative displacement to the target object.
[0127] The unavoidable zone of the target object refers to the area in which the virtual object cannot successfully avoid the target object when performing an avoidance action with a fixed action range. In other words, if the speech occurs when the virtual object is in the unavoidable zone, the virtual object cannot avoid colliding with the target object when performing an avoidance action with a fixed action range.
[0128] The invalid avoidance zone of the target object refers to the area where the virtual object has not yet reached the target object's location after performing an avoidance action with a fixed range of motion. In this case, the avoidance action is considered invalid.
[0129] Specifically, when voice is emitted in the video recording environment, the terminal responds by checking if the voice information matches the pre-configured interaction trigger information for the target object. It then determines whether the voice occurred when the virtual object was within an avoidable area of the target object in the video frame. If not, it further determines whether the voice occurred when the virtual object was within an unavoidable area of the target object in the video frame. If the voice occurred when the virtual object was within an unavoidable area of the target object in the video frame, the video recording screen displays a collision between the virtual object and the target object, showing the failed avoidance action of the virtual object towards the target object. If the voice did not occur when the virtual object was within an unavoidable area of the target object in the video frame, it is determined that the voice occurred when the virtual object was within an invalid avoidance area of the target object in the video frame. A preset fixed motion amplitude is then obtained, and the video recording screen displays the virtual object performing an avoidance action towards the target object according to the fixed motion amplitude, showing the invalid avoidance action of the virtual object continuing to move towards the target object.
[0130] refer to Figure 11The video recording shown uses a horizontal line segment from point A to point F to represent the horizontal movement path. The intersection of the virtual object's central axis and the horizontal line segment from point A to point F represents the virtual object's position. Within this horizontal line segment from point A to point F, the area between points C and D represents the non-contact, avoidable area of the target object; the areas between points B and C, and between points D and E, represent the contact, avoidable areas; the area between point E and point F represents the unavoidable area; and the area between point A and point B represents the ineffective avoidable area. If, during the virtual object's movement from point A to the target object, the user emits a matching voice command while the virtual object is within the area between points A and B, the video recording will show the virtual object performing an avoidance action against the target object with a fixed amplitude, demonstrating the ineffective avoidance effect of the virtual object continuing to move towards the target object. If, while the user emits a matching voice command while the virtual object is within the area between points B and C, the video recording will show the virtual object performing an avoidance action against the target object with a fixed amplitude, demonstrating the ineffective avoidance effect of the avoidance action. On the recording screen, the virtual object performs an avoidance maneuver against the target object with a fixed amplitude, demonstrating the effect of the virtual object successfully avoiding the target object with slight contact. If the user emits a matching voice message while the virtual object is in the area between points C and D, the recording screen shows the virtual object performing an avoidance maneuver against the target object with a fixed amplitude, demonstrating the effect of the virtual object successfully avoiding the target object without contact. If the user emits a matching voice message while the virtual object is in the area between points D and E, the recording screen shows the virtual object performing an avoidance maneuver against the target object with a fixed amplitude, demonstrating the effect of the virtual object successfully avoiding the target object with slight contact. If the user emits a matching voice message while the virtual object is in the area between points E and F, the recording screen shows the virtual object colliding with the target object, demonstrating the effect of the virtual object's failed avoidance maneuver against the target object.
[0131] In the above embodiments, the terminal sets up unavoidable and invalid avoidance areas corresponding to the target object, allowing the user to issue voice commands at appropriate times based on the relative positional relationship between the virtual object and the target object. This guides the virtual object to perform avoidance actions. By issuing voice commands in a timely and accurate manner, the user can effectively control the behavior of the virtual object to successfully avoid the target object. This interaction method not only enhances the user's participation and control capabilities but also improves the real-time performance and accuracy of the interaction. Furthermore, the avoidance result of the virtual object against the target object reflects the accuracy of the user's control over the timing of issuing voice commands, thereby increasing the user's interest in participating in video recording.
[0132] In one embodiment, the above information interaction processing method further includes the following steps: when the virtual object successfully avoids the target object, displaying the movement state of the virtual object relative to the other target objects.
[0133] Other target objects refer to any possible target objects other than those that have been successfully avoided. They can be the same as or different from the target objects that have been successfully avoided. For example, all the target objects can be railings, or the first target object can be a railing and the second target object can be a stone, etc.
[0134] The movement state refers to the change in position of a virtual object on the video recording screen.
[0135] Specifically, after the virtual object completes the avoidance action and successfully avoids the target object, the terminal can also determine the virtual object's next movement path and speed based on the virtual object's current position, preset path algorithm, and movement speed algorithm, and display the movement status of the virtual object moving towards other target objects according to the determined movement path and speed on the video recording screen.
[0136] In the above embodiments, by displaying the movement status of a virtual object moving towards other target objects after it successfully avoids a target object, the terminal can enhance the interactivity of the video, guide the user to pay attention to the next action of the virtual object, and enable the user to issue corresponding voice commands based on the movement status of the virtual object. This allows for control over the virtual object's avoidance of other target objects, improving the user's participation and control over the video content, thereby further enhancing the personalization of the video content. Simultaneously, the avoidance result of the virtual object towards the target object demonstrates the accuracy of the user's voice control over the virtual object's avoidance actions, thus increasing the user's interest in participating in video recording.
[0137] In one embodiment, when a virtual object successfully avoids a target object, the process of displaying the movement state of the virtual object relative to other target objects may specifically include the following steps: when the virtual object successfully avoids the target object in a non-contact manner, the video recording screen displays the movement state of the virtual object relative to other target objects at its original movement speed, where the size of other target objects is larger than the size of the target object, or the distance between other target objects is smaller than the distance between the target objects; or, when the virtual object successfully avoids the target object in a non-contact manner, the video recording screen displays the movement state of the virtual object relative to other target objects at a first movement speed, where the first movement speed is greater than the original movement speed.
[0138] The initial movement speed refers to the initial movement speed of the virtual object before it avoids the current target object. Specifically, it can be the speed at which the virtual object starts moving, or the previous movement speed during the avoidance process.
[0139] The size of other target objects is greater than the size of the target object, meaning that the size of other target objects displayed later is greater than the size of the target object that has been successfully avoided. The distance between other target objects is less than the distance between the target objects, meaning that the distance between other target objects displayed later and their previous target object is less than the distance between the target object that has been successfully avoided and its previous target object.
[0140] Specifically, after determining that a virtual object has successfully avoided a target object without contact, the terminal can increase the difficulty for subsequent virtual objects to avoid the target object. Specifically, it can increase the size of subsequent target objects and display the virtual object moving towards larger target objects at its original speed on the video recording screen; or it can decrease the distance between subsequent target objects and display the virtual object moving towards smaller target objects at its original speed on the video recording screen; or it can increase the speed of the virtual object and display the virtual object moving towards other target objects at a speed greater than its original speed at a distance that remains unchanged on the video recording screen.
[0141] In one embodiment, after determining that a virtual object has successfully avoided a current target object without contact, the terminal can further determine the number of consecutive times the virtual object has successfully avoided a target object without contact. When the number of contactless avoidances reaches a first threshold, the video recording screen displays the virtual object moving towards other target objects at its original speed, where the size of the other target objects is larger than the size of the target object, or the distance between other target objects is smaller than the distance between the target objects; or the video recording screen displays the virtual object moving towards other target objects at a first speed greater than the original speed. When the number of contactless avoidances does not reach the first threshold, the video recording screen displays the virtual object moving towards other target objects at its original speed, where the size of the other target objects is equal to the size of the target object, or the distance between other target objects is smaller than the distance between the target objects; or the video recording screen displays the virtual object moving towards other target objects at a first speed equal to the original speed.
[0142] In the above embodiments, after the terminal determines that the virtual object has successfully avoided the target object in a contactless manner, that is, when the user has performed good control over the virtual object to avoid the target object, the terminal can increase the difficulty of the virtual object avoiding the target object by increasing the size of the target object, decreasing the distance between the target objects, and increasing the movement speed of the virtual object. The increase in difficulty will stimulate the user's desire for challenge, thereby increasing the user's interest in participating in video recording.
[0143] In one embodiment, when a virtual object successfully avoids a target object, the process of displaying the movement state of the virtual object relative to other target objects may specifically include the following steps: when the virtual object successfully avoids the target object by contact, the video recording screen displays the movement state of the virtual object relative to other target objects at its original moving speed, where the size of the other target objects is less than or equal to the size of the target object, or the distance between other target objects is greater than or equal to the distance between the target objects; or, when the virtual object successfully avoids the target object by contact, the video recording screen displays the movement state of the virtual object relative to other target objects at a second moving speed, where the second moving speed is less than or equal to the original moving speed.
[0144] Among them, "the size of other target objects is less than or equal to the size of the target object" means that the size of other target objects displayed subsequently is less than or equal to the size of the target object that has been successfully avoided; "the distance between other target objects is less than the distance between target objects" means that the distance between other target objects displayed subsequently and their previous target object is less than the distance between the target object that has been successfully avoided and its previous target object.
[0145] Specifically, after determining that a virtual object has successfully avoided a target object through contact, the terminal can reduce the difficulty for subsequent virtual objects to avoid the target object. This can be achieved by reducing the size of subsequent target objects, displaying the virtual object moving at its original speed towards smaller target objects on the video recording screen; or by increasing the distance between subsequent target objects, displaying the virtual object moving at its original speed towards larger target objects on the video recording screen; or by decreasing the virtual object's speed, displaying it moving at a second speed (less than its original speed) towards target objects with the same distance on the video recording screen. Alternatively, after determining that a virtual object has successfully avoided a target object through contact, the terminal can also leave the difficulty for subsequent virtual objects to avoid the target object unchanged. Specifically, this can be achieved by displaying the virtual object moving at its original speed towards target objects of the same size, moving at its original speed towards target objects with the same distance on the video recording screen, or moving at its original speed towards target objects with the same distance on the video recording screen, or moving at its original speed towards target objects with the same distance on the video recording screen.
[0146] In one embodiment, after determining that a virtual object has successfully avoided a current target object through contact, the terminal can further determine the number of consecutive contactless avoidances the virtual object has made. When the number of contactless avoidances reaches a second threshold, the video recording screen displays the virtual object moving towards other target objects at its original speed, wherein the size of the other target objects is smaller than the size of the target object, or the distance between other target objects is greater than the distance between the target object; or the video recording screen displays the virtual object moving towards other target objects at a second speed, where the second speed is less than the original speed. When the number of contactless avoidances has not reached the second threshold, the video recording screen displays the virtual object moving towards other target objects at its original speed, wherein the size of the other target objects is equal to the size of the target object, or the distance between other target objects is less than the distance between the target object; or the video recording screen displays the virtual object moving towards other target objects at a second speed, where the second speed is equal to the original speed.
[0147] In the above embodiments, after determining that the virtual object has successfully avoided the target object in a contact manner, that is, when the user performs the avoidance action control on the virtual object, the terminal reduces the difficulty of the virtual object avoiding the target object in the subsequent process by reducing the size of the target object, increasing the distance between the target objects, and reducing the movement speed of the virtual object. This makes it easier for users to have a successful experience, enhances their sense of accomplishment and satisfaction, and thus increases the user's interest in participating in video recording.
[0148] In one embodiment, the above-described information interaction processing method further includes the following steps: when the voice information does not match the interaction trigger information pre-configured for the target object, displaying a first interaction prompt message for voice interaction.
[0149] The first interactive prompt information is used to prompt the user to emit a voice that matches the pre-configured interactive trigger information for the target object. Specifically, it can be at least one of text, icons, and image information, such as prompts related to increasing volume, clarity, or changing pronunciation. For example, if the user's voice volume is not greater than the pre-configured volume threshold for the target object, then the video recording screen will display "Shout out 'XX Milk' louder" or display a sound icon to indicate that the volume needs to be increased.
[0150] Specifically, when the terminal recognizes voice in the environment of the recorded video, it can extract voice from the recorded audio to obtain voice segments, perform information recognition on the voice segments to obtain voice information, and determine whether the obtained voice information matches the pre-configured interaction trigger information for the target object. If not, it displays the first interaction prompt information for voice interaction on the video recording screen, and simultaneously displays the virtual object moving towards the target object until it touches the target object. Alternatively, when voice is recognized again in the environment of the recorded video, the voice information is extracted, and the step of determining whether the obtained voice information matches the pre-configured interaction trigger information for the target object is repeated.
[0151] In the above embodiments, when the voice information does not match the pre-configured interaction trigger information for the target object, the terminal can provide timely feedback by displaying the first interaction prompt information, guiding the user to correctly adjust the voice input so as to effectively control the avoidance action of the virtual object, improve the effectiveness of the user's participation in the video content, and thus increase the user's interest in participating in video recording.
[0152] In one embodiment, the above information interaction processing method further includes the following steps: when the virtual object fails to avoid the target object, or when the virtual object successfully avoids all target objects, the target object avoidance score of the virtual object is displayed on the video recording screen; after the terminal obtains the video, it can also respond to the video sharing trigger operation and display video description information containing the target object avoidance score; and respond to the video sharing confirmation operation and share the video and video description information.
[0153] Among them, the avoidance score refers to the degree or performance of the virtual object in successfully avoiding the target object during the video recording process. It can be presented in the form of scores, grades or text descriptions so that users can understand how well they have performed in avoiding the target object.
[0154] The video sharing trigger action is the action used to trigger video sharing. Specifically, it can be a trigger action on the share button, selecting the share option, or other similar interface elements. The trigger action can be a click action.
[0155] The video sharing confirmation action is a confirmation action that needs to be performed after the user selects to share a video. This action can be triggered by a confirmation button, selecting a confirmation option, or other similar interface elements. The specific trigger action can be a click action.
[0156] Sharing videos and video descriptions can specifically involve sending the video and video description to social media friends or publishing the video and video description to social media platforms.
[0157] Specifically, each time a virtual object performs an avoidance action and successfully avoids a target object, the terminal analyzes the contact between the virtual object and the target object to obtain an avoidance evaluation result for each successful avoidance. When the virtual object fails to avoid the target object or successfully avoids all target objects, the terminal determines the virtual object's target object avoidance score based on the avoidance evaluation result for each successful avoidance and displays the virtual object's target object avoidance score on the video recording screen. After obtaining the video, the terminal can display the obtained video on a video preview page. The video preview page can display a sharing control. In response to the triggering operation of the sharing control, the terminal generates video description information based on the virtual object's target object avoidance score and displays the generated video description information and a confirmation control on the video preview page. In response to the triggering operation of the confirmation control, the terminal shares the video and the corresponding video description information.
[0158] In the above embodiments, the terminal can provide timely feedback to users by displaying the target object avoidance results of the virtual object, allowing users to understand their interactive performance. At the same time, in response to the trigger operation of the sharing control, video description information is generated based on the target object avoidance results of the virtual object. In response to the video sharing confirmation operation, the video and video description information are shared, allowing users to showcase their interactive performance, enhancing the interaction between different users, stimulating users' competitive desire, and further increasing users' interest in participating in video recording.
[0159] In one embodiment, the above information interaction processing method further includes the following steps: capturing a user's facial image during video recording; and displaying the user's facial image in the head area of the virtual object during virtual object display.
[0160] Among them, facial images refer to images of the user's face captured during video recording, which may include the user's facial features such as eyes, nose, and mouth.
[0161] Specifically, in response to the video recording operation, the terminal activates the camera device to capture the user's facial image in real time during the video recording process. While displaying the video recording screen in real time, the user's facial image is embedded into the head area of a virtual object to obtain a synthesized virtual object, which is then displayed in real time on the pre-recorded video screen.
[0162] like Figure 12 The image shown is a virtual object in one embodiment. The head region of the virtual object is blank. After embedding the user's facial image into the head region of the virtual object, the following can be obtained: Figure 13 The synthesized virtual object shown.
[0163] In the above embodiments, the terminal displays the user's facial image in the head area of the virtual object. The user can see their own facial expressions and changes in expression on the virtual object, making it easier for them to associate the virtual object with themselves. This increases the fun and appeal of the video content, enhances the user's interest in participating in video recording, and also enables personalized customization of the video content, further improving the personalization of the video content.
[0164] In one embodiment, the above information interaction processing method further includes the following steps: displaying a video recording page; displaying candidate recording templates on the video recording page; in response to a template selection operation, displaying a second interactive prompt message for the target recording template specified by the template selection operation; and in response to a video recording start operation, displaying a video recording screen for the target recording template, so that the user can perform interactive actions on the target object by controlling the virtual object through voice during the video recording process.
[0165] The video recording page is used to configure various parameters and functions for video recording, such as recording duration and recording template.
[0166] Candidate recording templates are preset templates or styles that users can choose to record videos. These templates usually include specific video styles, layouts, effects, etc. Users can choose the corresponding template to record videos according to their needs and preferences. Selecting a template can help users create video content with specific styles and effects more quickly.
[0167] The second interactive prompt is used to guide users on things to pay attention to during video recording, or to provide voice interaction operation guidance related to the target recording template. Specifically, it can be text, icons, or other forms.
[0168] Specifically, in response to the content generation operation, the terminal displays a video recording page, which shows at least one candidate recording template. The user can select from the displayed candidate recording template. In response to the template selection operation, the terminal displays a video recording preview and recording controls corresponding to the target recording template specified by the template selection operation on the video recording page. The video recording preview displays a second interactive prompt. In response to the triggering operation of the recording controls, the video recording screen for the target recording template is displayed on the video recording page, so that the user can perform interactive actions on the target object through voice control of the virtual object during the video recording process.
[0169] The video recording preview screen refers to the preview screen displayed on the terminal when the user selects a recording template and is ready to start recording a video. This screen can display the position of the virtual object, scene settings, and secondary interactive prompts. Users can preview the interactive effects of voice-controlled virtual objects on this screen so as to make necessary adjustments or preparations.
[0170] like Figure 6 The video recording page shown displays a recording control 602 and multiple candidate recording templates, such as Template 1, Template 2, Template 3, and Template 4. When the user selects Template 3, a preview screen 604 is displayed, showing the second prompt message "Ready! Go! Shout 'I love sports' to complete the hurdles." When the user clicks the recording control 602, the terminal responds to the click operation and begins recording the video. Figure 7 As shown, the video recording screen 702 is displayed on the video recording page at the same time. According to the second interactive prompt information, the user can use voice to control the virtual object to perform avoidance actions against the target object during the video recording process.
[0171] In the above embodiments, by displaying a video recording page and providing candidate recording templates, the terminal allows users to more conveniently select a suitable recording template. Users can choose the template they want to use from the options. By displaying the second interactive prompt information of the target recording template, the terminal can guide users to better understand the characteristics and usage of the recording template, improve the accuracy of users' control over virtual objects to perform avoidance actions against target objects through voice during video recording, and thus increase users' interest in participating in video recording.
[0172] This application also provides an information interaction processing method, which is executed by a computer device, such as a computer device that may be... Figure 1 The terminal 102 or server 104 shown is used by this interaction method. Figure 1 The following explanation uses the terminal 102 shown as an example. Figure 14 As shown, the method may include the following steps:
[0173] S1402, Display a virtual object that has a relative displacement with respect to the target object.
[0174] Virtual objects are graphics, models, or entities created in a digital environment. They can be generated through computer graphics technology and can take various forms, such as virtual characters or virtual animals.
[0175] The target object is an element used to obstruct the movement of a virtual object. It can be generated through computer graphics technology and can take various forms, such as virtual vehicles, virtual railings, virtual animals, virtual characters, virtual trees, virtual ditches, etc.
[0176] In one embodiment, the terminal interaction is initiated by displaying an interactive screen, which shows a virtual object that has a relative displacement with the target object according to a preset initial moving speed. After the virtual object successfully interacts with the target object, the interaction displays the movement state of the virtual object with an adjusted moving speed relative to other target objects, wherein the adjusted moving speed may be the same as or different from the initial moving speed.
[0177] S1404, Display recommended information containing prompts for voice input.
[0178] The prompts are used to guide the input of voice commands, while the recommendations are used to recommend specific products, services, or brands. For example, a recommended message might be "XX Milk is so delicious, shout 'XX Milk' to complete the hurdles." Here, XX Milk is a milk brand, and the recommendation is specifically for XX Milk. The prompt "Shout 'XX Milk' to complete the hurdles" is the information included in this recommendation.
[0179] S1406, In response to receiving the voice, display the virtual object to perform an interactive action on the target object.
[0180] Speech refers to sound signals produced through vocal organs (such as the mouth and vocal cords) for communication and information transmission; specifically, it can be speaking or singing.
[0181] Interactive actions are the responses of virtual objects to voice commands input by users. Specifically, interactive actions can be movement actions targeting a target object, such as moving to a specific location, or avoidance actions. Avoidance actions aim to avoid target objects that hinder the movement of virtual objects, such as jumping upwards, moving to the left, or moving to the right.
[0182] In one embodiment, during the interaction, the terminal can record sounds in the environment using an audio recording device to obtain recorded audio, and perform speech recognition on the recorded audio in real time. When speech is detected in the environment of the recorded video, speech can be extracted from the recorded audio to obtain speech segments, and information recognition can be performed on the speech segments to obtain speech information. It is then determined whether the obtained speech information matches the pre-configured interaction trigger information for the target object. If it matches, a virtual object is displayed on the interaction screen to perform an interaction action for the target object. If it does not match, the virtual object is displayed on the interaction screen to maintain a movement state with relative displacement to the target object until it collides with the target object. Alternatively, when speech is detected again in the environment of the recorded video, the speech information is extracted, and the step of determining whether the obtained speech information matches the pre-configured interaction trigger information for the target object is re-executed.
[0183] Speech information refers to specific aspects of speech characteristics, which can be at least one of content information and speech attribute information. Content information, also known as semantic information, refers to the actual meaning or connotation conveyed in speech, i.e., text content. Speech attribute information refers to the characteristics and properties of the sound itself, which can be at least one of speech volume, speech pitch, and speech rhythm. These attributes reflect the characteristics of the speaker's sound wave amplitude, frequency changes, and rhythmic patterns during the process of vocalization.
[0184] Interactive trigger information is pre-configured information used to trigger virtual objects to perform specific actions. Specifically, it can be at least one of the following sub-voice trigger information: volume trigger information, tone trigger information, and content trigger information. Specifically, volume trigger information can be a volume trigger threshold, tone trigger information can be a preset voice tone, and content trigger information can be preset voice content. Interactive trigger information can also be song trigger information from song rhythm trigger information, song melody trigger information, and song content trigger information. Specifically, song rhythm trigger information can be the rhythm information of the corresponding song segment of the target song, song melody trigger information can be the melody information of the corresponding song segment of the target song, and song content trigger information can be the lyrics information of the corresponding song segment of the target song.
[0185] Matching conditions are used to determine whether the actual voice input received in the recording video environment matches the preset interaction trigger information. For example, the content trigger information might be keywords like "left," "right," or "jump." In a real-world scenario, when the similarity between the user's spoken words and a preset keyword reaches a content similarity trigger threshold, the voice information is determined to match the pre-configured interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. Conversely, if the similarity between the user's spoken words and any preset keyword does not reach the similarity trigger threshold, the voice information is determined to not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action. Another example is the volume trigger information, which is a certain volume threshold. In a real-world scenario, when the user's voice volume reaches this threshold, the voice input is determined to match the preset interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. If the information matches the pre-configured interaction trigger information for the target object, the virtual object can be triggered to perform the corresponding action. If the volume of the user's voice does not reach the volume threshold, it is determined that the voice information does not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action. For example, if the tone trigger information is a preset voice tone, if the similarity between the user's voice tone and the preset voice tone reaches the tone similarity trigger threshold, it is determined that the voice information matches the pre-configured interaction trigger information for the target object, and the virtual object can be triggered to perform the corresponding action. If the similarity between the user's voice tone and the preset voice tone does not reach the tone similarity trigger threshold, it is determined that the voice information does not match the pre-configured interaction trigger information for the target object, and the user will not be triggered to perform the corresponding action.
[0186] S1408, in response to the trigger operation of the recommendation information, displays the recommendation information interface.
[0187] Among them, the triggering operation for recommendation information is the operation by which the user displays more information associated with the recommendation information. Specifically, it can be a click operation, a hover trigger operation, a sound activation operation, a scan operation, a swipe operation, etc.
[0188] The recommendation information interface is a page used to display more information associated with the recommended information. Specifically, it can be a pop-up window, a side menu, a new page layer, or a modal dialog box. For example, when a user clicks on the recommendation "XX Milk is so delicious, shout 'XX Milk' to complete the hurdle," a product information page for "XX Milk" is displayed. This product information page shows product images, product information, and purchase links for "XX Milk."
[0189] Specifically, in response to the trigger operation of the displayed recommendation information, the terminal obtains relevant recommendation data, generates a recommendation information interface based on the recommendation data, and displays the generated recommendation information interface.
[0190] In the above embodiments, the terminal displays an interactive virtual object with relative displacement to the target object. The user can see the positional changes of the virtual object relative to the target object. By displaying recommendation information containing prompts, the user can be prompted to input voice to control the interaction between the virtual object and the target object. In response to receiving voice input, the virtual object is displayed to perform interactive actions against the target object. In response to triggering the recommendation information, the recommendation information interface is displayed. This allows the user to interact with the virtual object and the target object through voice control, and to interact with the recommendation information to obtain more information while controlling the virtual object. This avoids the situation where the user can only switch to other interfaces to obtain more information related to the recommendation information after the interaction between the virtual object and the target object is completed, thus improving the efficiency of information acquisition.
[0191] In one embodiment, the process of a terminal displaying a recommendation information interface in response to a triggering operation on recommendation information includes the following steps: in response to a triggering operation on recommendation information, if the virtual object successfully avoids the target object, a first recommendation information interface is displayed; if the virtual object fails to successfully avoid the target object, a second recommendation information interface is displayed.
[0192] The first recommendation information interface displays more information associated with the recommendation information and first additional information. The second recommendation information interface displays more information associated with the recommendation information and second additional information. The first additional information and the second additional information are respectively related to the virtual object's avoidance result against the target object. Specifically, the first additional information may be related to the virtual object's successful avoidance of the target object, and the second additional information may be related to the virtual object's unsuccessful avoidance of the target object. For example, the first additional information may be reward information for successful avoidance, and the reward information may specifically be discount information for the products displayed in the recommendation information interface. The second additional information may be analysis information on the reasons for failed avoidance, etc.
[0193] Specifically, in response to the triggered operation of the displayed recommendation information, the terminal obtains relevant recommendation data and determines the avoidance result of the virtual object towards the target object. When the avoidance result is that the virtual object successfully avoids the target object, the terminal determines the corresponding first additional information, generates a first recommendation information interface based on the recommendation data and the first additional information, and displays the generated first recommendation information interface. When the avoidance result is that the virtual object fails to avoid the target object, the terminal determines the corresponding second additional information, generates a second recommendation information interface based on the recommendation data and the second additional information, and displays the generated second recommendation information interface.
[0194] In the above embodiments, the terminal responds to the triggering operation of the recommendation information. If the virtual object successfully avoids the target object, a first recommendation information interface is displayed; if the virtual object fails to avoid the target object, a second recommendation information interface is displayed. By distinguishing between successful and unsuccessful avoidance, specific and personalized feedback can be provided to the user. At the same time, it avoids the situation where the user can only switch to other interfaces to obtain more information related to the recommendation information and interaction results after the virtual object has interacted with the target object, thus improving the efficiency of information acquisition.
[0195] In one embodiment, the process of a terminal displaying a virtual object performing an interactive action against a target object includes the following steps: displaying the virtual object to perform an avoidance action against the target object according to the action amplitude associated with the voice information; wherein, when the action amplitude meets a preset amplitude condition, the virtual object successfully avoids the target object.
[0196] Among them, the range of motion refers to the degree or range exhibited by the virtual object when performing an avoidance action. For example, when the avoidance action is an upward hurdle jump, the corresponding range of motion can be the height of the jump or the vertical distance.
[0197] The specific range of motion associated with voice information can be adjusted based on the voice information itself. Voice information can include at least one sub-voice information such as voice volume, voice tone, or voice content. Voice information can also include at least one song information such as the rhythm, melody, or content of a sung song. For example, when the voice information is voice volume, the corresponding motion range can be adjusted based on the volume level. When the voice information is voice content, the corresponding motion range can be adjusted based on the similarity between the voice content and the content trigger information. When the voice information is voice tone, the corresponding motion range can be adjusted based on the similarity between the voice tone and a preset voice tone. When the voice information is the rhythm of a sung song, the corresponding motion range can be adjusted based on the similarity between the rhythm of the sung song and the standard rhythm of the target song. When the voice information is the melody of a sung song, the corresponding motion range can be adjusted based on the similarity between the melody of the sung song and the standard melody of the target song. When the voice information is the content of a sung song, the corresponding motion range can be adjusted based on the similarity between the content of the sung song and the standard content of the target song. Here, "sung song" refers to the version of the song expressed by the user through singing or chanting the target song, and "target song" refers to the original recorded version of the song, also known as the original version.
[0198] The preset amplitude condition is a pre-defined range of motion that a virtual object must achieve to successfully avoid a target object. In other words, if the range of motion of the virtual object when performing the avoidance action meets the preset amplitude condition, then the virtual object can successfully avoid the target object, that is, the avoidance action is successful. Conversely, if the range of motion does not meet the preset amplitude condition, that is, it does not reach the required level or amplitude, then the virtual object's avoidance action fails, which may result in a collision or occlusion with the target object.
[0199] The motion amplitude condition can be specifically a motion amplitude threshold condition. A threshold or range can be set. When the motion amplitude of the virtual object reaches the threshold or is within the range, the motion amplitude of the virtual object is determined to meet the preset amplitude condition.
[0200] Specifically, after receiving voice information, the terminal can determine the corresponding action amplitude of the virtual object based on the voice information and the amplitude function corresponding to the voice information. The virtual object is then displayed on the video recording screen to perform an avoidance action against the target object according to the determined action amplitude. For example, the motion trajectory and posture of the virtual object can be determined based on the determined action amplitude, the distance between the virtual object and the target object, and the movement speed of the virtual object. The motion trajectory and posture of the virtual object are displayed on the video recording screen to show the execution effect of the avoidance action. When the action amplitude of the virtual object meets the preset amplitude conditions, the execution effect of the avoidance action of the virtual object successfully avoiding the target object is shown. When the action amplitude of the virtual object does not meet the preset amplitude conditions, the execution effect of the avoidance action of the virtual object colliding with the target object is shown.
[0201] In the above embodiments, the terminal controls the virtual object to perform avoidance actions against the target object and the range of motion based on the avoidance actions through voice information, thereby improving the user's control over the virtual object.
[0202] In one embodiment, the process of a terminal displaying a virtual object performing an interactive action against a target object includes the following steps: displaying the virtual object to perform an avoidance action against the target object according to a fixed action range; wherein, if the voice occurs when the virtual object is in the avoidable area of the target object in the video frame, the virtual object successfully avoids the target object.
[0203] Among them, the fixed motion amplitude is a pre-set, fixed and unchanging degree or range exhibited by the virtual object when performing an avoidance action, and its size is not affected by voice information.
[0204] The avoidable area of the target object refers to the area in which the virtual object can successfully avoid the target object when performing an avoidance action with a fixed action range. It can be preset according to the shape, size, fixed action range, etc. of the target object.
[0205] Specifically, when voice is emitted in the video recording environment, the terminal responds to the voice information and determines whether the voice occurs when the virtual object is in the avoidable area of the target object in the video frame, if the voice information matches the pre-configured interaction trigger information for the target object. If so, the terminal obtains the preset fixed motion amplitude and displays the virtual object performing the avoidance action against the target object according to the fixed motion amplitude on the video recording screen, presenting the effect of the virtual object successfully avoiding the target object.
[0206] In the above embodiments, the terminal sets a dodgeable area corresponding to the target object, allowing the user to issue voice commands at appropriate times based on the relative positional relationship between the virtual object and the target object. This guides the virtual object to perform dodge actions. By issuing voice commands in a timely and accurate manner, the user can effectively control the behavior of the virtual object to successfully avoid the target object. This interaction method not only enhances the user's participation and control capabilities but also improves the real-time performance and accuracy of the interaction. Furthermore, the successful avoidance of the target object by the virtual object demonstrates the accuracy of the user's control over the timing of issuing voice commands, thereby increasing the user's interest in controlling the virtual object.
[0207] This application also provides an application scenario, specifically a video shooting scenario on a short video platform. The aforementioned information interaction processing method can be applied to this scenario and implemented through an information interaction processing system. The architecture of this system is a "client-server" structure, also known as a "terminal-server" structure. The terminal runs a short video shooting module, a game template module, and a short video publishing module. The server stores relevant resources for the short video game templates. When a user triggers a short video shooting operation on the terminal, the terminal responds by activating the game template module, thereby dynamically retrieving the corresponding template resources from the server, uploading the completed short video, and performing subsequent post-processing such as sharing and distribution. (See reference...) Figure 15 The interaction sequence diagram shown illustrates that the interactive system implementing the aforementioned information interaction processing method may specifically include the following steps:
[0208] I. Before starting filming.
[0209] Terminal display as follows Figure 6 On the video recording page shown, users can select a target game template. In response to the user's selection, the terminal activates the game template module to obtain the resources corresponding to the target game template from the server. After obtaining the resources corresponding to the target game template, the terminal calls the resources corresponding to the target game template through the video recording template.
[0210] Simultaneously, the short video shooting template is activated, and the camera is started through the short video shooting module to capture the user's facial image. The game template module fills the facial image into the head area of the virtual object and synthesizes the game running screen. The synthesized game running screen data is then sent to the shooting module frame by frame according to the frame rate setting to be displayed to the user.
[0211] When video recording begins, which is also when the game starts, the game module will display the initial screen to guide the user through the basic rules of the game, such as the words that need to be spoken.
[0212] 2. After the user clicks to take a picture.
[0213] like Figure 6 As shown, users can click the recording control on the video recording page. The terminal responds to the trigger operation of the recording control, starts the game module, and the game module starts running and calls the microphone to obtain ambient sound in real time.
[0214] In addition, the game module runs a speech recognition algorithm locally, such as a deep learning model trained based on Mel frequency cepstral coefficient (MFCC) features [1]. When the acquired speech meets the preset matching conditions, it controls the virtual object to perform the target object avoidance action. For example, when the recognition confidence is greater than a certain preset threshold, it drives the user character to perform a jumping action.
[0215] Simultaneously, based on a preset algorithm, the collision between the virtual object and the target object is determined, and the process of the virtual object performing the avoidance action and the avoidance result are displayed on the video recording page.
[0216] III. When the game process ends.
[0217] When the terminal detects that a game has ended, the game module sends a message to the camera module to stop recording.
[0218] The shooting module converts and processes the collected video data, and the publishing module saves the video to the server. The server returns the saving result to the terminal, and the terminal can display the obtained video to the user.
[0219] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0220] Based on the same inventive concept, this application also provides an information interaction processing apparatus for implementing the information interaction processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more information interaction processing apparatus embodiments provided below can be found in the limitations of the information interaction processing method described above, and will not be repeated here.
[0221] In one embodiment, such as Figure 16 As shown, an information interaction processing device is provided, including: a virtual object display module 1602, a recommendation information display module 1604, a voice control module 1606, and a video acquisition module 1608, wherein:
[0222] The virtual object display module 1602 is used to display virtual objects that have a relative displacement with respect to the target object;
[0223] The recommendation information display module 1604 is used to display recommendation information containing prompts, which are used to prompt for voice input.
[0224] The voice control module 1606 is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object;
[0225] The video acquisition module 1608 is used to acquire the video in response to the end of video recording. The video contains the content displayed on the video recording screen during the video recording process.
[0226] In the above embodiments, during video recording, by displaying an interactive virtual object with relative displacement to the target object, the user can see the positional changes of the virtual object relative to the target object. By displaying recommended information containing prompts, the user can be prompted to input voice to control the interaction between the virtual object and the target object. In response to receiving voice input, the virtual object is displayed to perform interactive actions against the target object. This allows the user to interact with the virtual object and the target object via voice control during video recording. At the end of video recording, a video recording the interaction between the virtual object and the target object based on voice is obtained. By allowing users to use voice control to interact with the virtual object and the target object during video recording, videos generated by different users are unique, and the content of each video recorded by the same user is also different, enhancing the user's participation and control over the video content, thereby further improving the personalization of the video content. At the same time, the interaction results of the virtual object with the target object can also demonstrate the accuracy of the user's control over the virtual object's interactive actions via voice, thereby increasing the user's interest in participating in video recording.
[0227] In one embodiment, the voice control module 1606 is further configured to: display on the video recording screen that the virtual object performs an avoidance action against the target object according to the action amplitude associated with the voice information; wherein, when the action amplitude meets a preset amplitude condition, the virtual object successfully avoids the target object.
[0228] In one embodiment, the voice information includes at least one sub-voice information among voice volume, voice tone, or voice content, and the action amplitude includes a first action amplitude corresponding to at least one sub-voice information; the voice control module 1606 is further configured to: display on the video recording screen that a virtual object performs an avoidance action against a target object according to the first action amplitude.
[0229] In one embodiment, the voice is a song sung by the user according to the target song, and the voice information includes at least one of the song information such as the song rhythm, song melody or song content. The action range includes a second action range corresponding to at least one song information. The voice control module 1606 is also used to: display on the video recording screen that the virtual object performs an avoidance action against the target object according to the second action range.
[0230] In one embodiment, such as Figure 17 As shown, the device also includes an accompaniment playback module 1610, used for: playing the accompaniment of the target song during video recording; and, in response to acquiring voice, displaying on the video recording screen the similarity between at least one song information and the standard song information of the target song.
[0231] In one embodiment, the preset amplitude conditions include a first amplitude condition and a second amplitude condition. The voice control module 1606 is further configured to: when the amplitude of the action meets the first amplitude condition, the virtual object successfully avoids the target object in a non-contact manner; and when the amplitude of the action meets the second amplitude condition, the virtual object successfully avoids the target object in a contact manner.
[0232] In one embodiment, the voice control module 1606 is further configured to: display on the video recording screen that a virtual object performs an avoidance action against a target object according to a fixed movement amplitude; wherein, if the voice occurs when the virtual object is in the avoidable area of the target object in the video screen, the virtual object successfully avoids the target object.
[0233] In one embodiment, the avoidable area includes a first avoidable area and a second avoidable area. The voice control module 1606 is further configured to: if the voice occurs when the virtual object is in the first avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in a non-contact manner; if the voice occurs when the virtual object is in the second avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in a contact manner.
[0234] In one embodiment, the virtual object display module 1602 is further configured to: if the voice occurs when the virtual object is in an unavoidable area of the target object in the video frame, the virtual object collides with the target object; if the voice occurs when the virtual object is in an invalid avoidance area of the target object in the video frame, after performing the avoidance action, the virtual object maintains a moving state with relative displacement to the target object.
[0235] In one embodiment, the virtual object display module 1602 is further configured to: display the movement state of the virtual object relative to other target objects when the virtual object successfully avoids the target object.
[0236] In one embodiment, the virtual object display module 1602 is further configured to: when the virtual object successfully avoids the target object in a non-contact manner, display on the video recording screen the movement state of the virtual object relative to other target objects at its original moving speed, wherein the size of other target objects is larger than the size of the target object, or the target object spacing corresponding to other target objects is smaller than the target object spacing corresponding to the target object; or, when the virtual object successfully avoids the target object in a non-contact manner, display on the video recording screen the movement state of the virtual object relative to other target objects at a first moving speed, wherein the first moving speed is greater than the original moving speed.
[0237] In one embodiment, the virtual object display module 1602 is further configured to: when the virtual object successfully avoids the target object by contact, display on the video recording screen the movement state of the virtual object relative to other target objects at its original moving speed, wherein the size of other target objects is less than or equal to the size of the target object, or the target object spacing corresponding to other target objects is greater than or equal to the target object spacing corresponding to the target object; or, when the virtual object successfully avoids the target object by contact, display on the video recording screen the movement state of the virtual object relative to other target objects at a second moving speed, wherein the second moving speed is less than or equal to the original moving speed.
[0238] In one embodiment, such as Figure 17 As shown, the device also includes an interactive prompt module 1612, used to: display first interactive prompt information when the voice information of the voice does not match the interaction trigger information pre-configured for the target object; the first interactive prompt information is used to guide the user to emit voice that matches the interaction trigger information pre-configured for the target object.
[0239] In one embodiment, such as Figure 17As shown, the device also includes a performance display module 1614, used to: display the virtual object's target object avoidance performance when the virtual object fails to avoid the target object or when the virtual object successfully avoids all target objects; and a video sharing module 1616, used to: display video description information containing the target object avoidance performance in response to a video sharing trigger operation; and share the video and video description information in response to a video sharing confirmation operation.
[0240] In one embodiment, the virtual object display module 1602 is further configured to: capture a user's facial image during video recording; and display the user's facial image in the head area of the virtual object during the display of the virtual object.
[0241] In one embodiment, such as Figure 17 As shown, the device also includes a recording configuration module 1618, used for: displaying a video recording page; displaying candidate recording templates on the video recording page; responding to a template selection operation, displaying a second interactive prompt message for the target recording template specified by the template selection operation; and responding to a video recording start operation, displaying a video recording screen for the target recording template, so that the user can perform interactive actions on the target object by controlling the virtual object through voice during the video recording process.
[0242] In one embodiment, such as Figure 18 As shown, an information interaction processing device is also provided, including: a virtual object display module 1802, a recommendation information display module 1804, a voice control module 1806, and a recommendation interface display module 1808, wherein:
[0243] The virtual object display module 1802 is used to display virtual objects that have a relative displacement with respect to the target object;
[0244] The recommendation information display module 1804 is used to display recommendation information containing prompts, which are used to prompt for voice input.
[0245] The voice control module 1806 is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object;
[0246] The recommendation interface display module 1808 is used to display the recommendation information interface in response to the trigger operation of recommendation information.
[0247] In the above embodiments, by displaying an interactive virtual object with relative displacement to the target object, the user can see the positional changes of the virtual object relative to the target object. By displaying recommended information containing prompts, the user can be prompted to input voice to control the interaction between the virtual object and the target object. In response to receiving voice input, the virtual object is displayed to perform interactive actions against the target object. In response to triggering the recommended information, the recommended information interface is displayed. This allows the user to interact with the virtual object and the target object through voice control, and to interact with the recommended information to obtain more information while controlling the virtual object. This avoids the situation where the user can only switch to other interfaces to obtain more information related to the recommended information after the interaction between the virtual object and the target object is completed, thus improving the efficiency of information acquisition.
[0248] In one embodiment, the recommendation interface display module 1808 is further configured to: in response to a triggering operation on the recommendation information, if the virtual object successfully avoids the target object, display a first recommendation information interface; if the virtual object fails to avoid the target object, display a second recommendation information interface.
[0249] In one embodiment, the voice control module 1806 is further configured to: display the virtual object performing an avoidance action against the target object according to the action amplitude associated with the voice information; wherein, when the action amplitude meets a preset amplitude condition, the virtual object successfully avoids the target object.
[0250] In one embodiment, the voice control module 1806 is further configured to: display the virtual object performing an avoidance action against the target object according to a fixed movement amplitude; wherein, if the voice occurs when the virtual object is in the avoidable area of the target object in the video frame, the virtual object successfully avoids the target object.
[0251] Each module in the aforementioned information interaction processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0252] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 19As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an information interaction processing method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0253] Those skilled in the art will understand that Figure 19 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0254] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0255] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0256] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0257] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0258] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0259] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0260] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An information interaction processing method, characterized in that, The method includes: During video recording, a virtual object with a relative displacement to the target object is displayed; Display recommended information containing prompts, which are used to prompt for voice input; In response to receiving voice, the virtual object is displayed to perform an interactive action on the target object; In response to the end of video recording, a video is obtained, the video containing the content displayed during the video recording process.
2. The method according to claim 1, characterized in that, The step of displaying the virtual object to perform interactive actions on the target object includes: The video recording shows the virtual object performing an avoidance maneuver against the target object according to the movement amplitude associated with the voice information. Specifically, when the range of motion meets a preset range condition, the virtual object successfully avoids the target object.
3. The method according to claim 2, characterized in that, The voice information includes at least one sub-voice information among voice volume, voice tone, or voice content, and the movement amplitude includes a first movement amplitude corresponding to the at least one sub-voice information; The step of displaying the virtual object on the video recording screen to perform an avoidance action against the target object according to the movement amplitude associated with the voice information includes: The video recording shows the virtual object performing an avoidance maneuver against the target object according to the first movement amplitude.
4. The method according to claim 2, characterized in that, The voice is a song sung by the user based on the target song. The voice information includes at least one of the song information such as the song rhythm, song melody, or song content. The movement amplitude includes a second movement amplitude corresponding to the at least one song information. The step of displaying the virtual object on the video recording screen to perform an avoidance action against the target object according to the movement amplitude associated with the voice information includes: The video recording shows the virtual object performing an avoidance maneuver against the target object according to the second action range.
5. The method according to claim 4, characterized in that, The method further includes: Play the instrumental version of the target song during video recording; In response to acquiring voice, the similarity between the at least one song information and the standard song information of the target song is displayed on the video recording screen.
6. The method according to claim 2, characterized in that, The preset amplitude conditions include a first amplitude condition and a second amplitude condition. When the motion amplitude meets the preset amplitude conditions, the virtual object successfully avoids the target object, including: When the range of motion meets the first range condition, the virtual object successfully avoids the target object in a non-contact manner; When the magnitude of the action meets the second magnitude condition, the virtual object successfully avoids the target object by making contact.
7. The method according to claim 1, characterized in that, The step of displaying the virtual object to perform interactive actions on the target object includes: The video recording shows the virtual object performing an avoidance maneuver against the target object with a fixed range of motion. If the voice occurs when the virtual object is within the avoidable area of the target object in the video frame, the virtual object successfully avoids the target object.
8. The method according to claim 7, characterized in that, The avoidable area includes a first avoidable area and a second avoidable area. If the voice occurs when the virtual object is within the avoidable area of the target object in the video frame, and the virtual object successfully avoids the target object, this includes: If the voice occurs when the virtual object is in the first avoidable area of the target object in the video frame, the virtual object successfully avoids the target object in a non-contact manner; If the voice occurs when the virtual object is in the second avoidable area of the target object in the video frame, the virtual object successfully avoids the target object by making contact.
9. The method according to claim 7, characterized in that, The method further includes: If the voice occurs when the virtual object is in the unavoidable area of the target object in the video frame, the virtual object collides with the target object. If the voice occurs when the virtual object is in the invalid avoidance zone of the target object in the video frame, after the avoidance action is completed, the virtual object maintains a movement state with relative displacement to the target object.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: When the virtual object successfully avoids the target object, the movement state of the virtual object relative to the other target objects is displayed.
11. The method according to claim 10, characterized in that, When the virtual object successfully avoids the target object, the display of the virtual object's movement state relative to other target objects includes: When the virtual object successfully avoids the target object without contact, the video recording shows the virtual object moving relative to other target objects at its original speed, where the size of the other target objects is larger than the size of the target object, or the distance between the other target objects is smaller than the distance between the target object; or... When the virtual object successfully avoids the target object in a non-contact manner, the video recording screen displays the virtual object moving relative to the other target objects at a first moving speed, where the first moving speed is greater than the original moving speed.
12. The method according to claim 10, characterized in that, When the virtual object successfully avoids the target object, the display of the virtual object's movement state relative to other target objects includes: When the virtual object successfully avoids the target object through contact, the video recording displays the virtual object's movement state relative to other target objects at its original moving speed, where the size of the other target objects is less than or equal to the size of the target object, or the distance between the other target objects is greater than or equal to the distance between the target objects; or... When the virtual object successfully avoids the target object by contact, the video recording screen displays the virtual object moving relative to the other target objects at a second moving speed, where the second moving speed is less than or equal to the original moving speed.
13. The method according to any one of claims 1 to 12, characterized in that, The method further includes: When the voice information does not match the pre-configured interaction trigger information for the target object, a first interaction prompt is displayed; the first interaction prompt is used to guide the user to emit voice that matches the pre-configured interaction trigger information for the target object.
14. The method according to any one of claims 1 to 12, characterized in that, The method further includes: When the virtual object fails to avoid the target object, or when the virtual object successfully avoids all target objects, the target object avoidance score of the virtual object is displayed. After obtaining the video, the method further includes: In response to a video sharing trigger, display video description information including the target object avoidance results; In response to the video sharing confirmation operation, the video and the video description information are shared.
15. The method according to any one of claims 1 to 12, characterized in that, The method further includes: During video recording, the user's facial image is captured; During the display of the virtual object, the user's facial image is displayed in the head area of the virtual object.
16. The method according to any one of claims 1 to 12, characterized in that, The method further includes: The video recording page is displayed; candidate recording templates are shown on the video recording page. In response to a template selection operation, a second interactive prompt message for the target recording template specified by the template selection operation is displayed; In response to the start of video recording, a video recording screen for the target recording template is displayed, allowing the user to control the virtual object to perform interactive actions on the target object via voice during the video recording process.
17. An information interaction processing method, characterized in that, The method includes: Displays a virtual object that has a relative displacement with respect to the target object; Display recommended information containing prompts, which are used to prompt for voice input; In response to receiving voice, the virtual object is displayed to perform an interactive action on the target object; In response to the triggering operation of the recommended information, the recommended information interface is displayed.
18. The method according to claim 17, characterized in that, The step of displaying the recommendation information interface in response to a trigger operation on the recommendation information includes: In response to the triggering operation of the recommendation information, if the virtual object successfully avoids the target object, a first recommendation information interface is displayed; if the virtual object fails to avoid the target object, a second recommendation information interface is displayed.
19. The method according to claim 17, characterized in that, The step of displaying the virtual object to perform interactive actions on the target object includes: The virtual object is shown to perform an avoidance action against the target object according to the movement amplitude associated with the voice information of the speech; Specifically, when the range of motion meets a preset range condition, the virtual object successfully avoids the target object.
20. The method according to claim 17, characterized in that, The step of displaying the virtual object to perform interactive actions on the target object includes: The virtual object is displayed to perform an avoidance action against the target object according to a fixed movement range; If the voice occurs when the virtual object is within the avoidable area of the target object in the video frame, the virtual object successfully avoids the target object.
21. An information interaction processing device, characterized in that, The device includes: The virtual object display module is used to display virtual objects that have a relative displacement to the target object during video recording. The recommendation information display module is used to display recommendation information containing prompts, the prompts being used to prompt for voice input; The voice control module is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object; The video acquisition module is used to acquire a video in response to the end of video recording, the video containing the content displayed on the video recording screen during the video recording process.
22. An information interaction processing device, characterized in that, The device includes: The virtual object display module is used to display virtual objects that have a relative displacement with respect to the target object; The recommendation information display module is used to display recommendation information containing prompts, the prompts being used to prompt for voice input; The voice control module is used to respond to the acquisition of voice and display the virtual object to perform interactive actions on the target object; The recommendation interface display module is used to display the recommendation information interface in response to a trigger operation on the recommendation information.
23. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 20.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 20.
25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 20.