Video stream generation method, electronic device and computer-readable storage medium
By adjusting the virtual camera settings and converting the virtual video stream using the target video generation model, the problem of insufficient realism of the virtual video stream is solved, and realistic simulation video stream generation is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- PCT/IB2024/062687
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-30
- Filing Date
- 2024-12-16
- Publication Date
- 2025-10-09
AI Technical Summary
Existing virtual video stream generation methods are insufficient in terms of picture realism and have limited application scenarios. In particular, the CARLA simulator is only suitable for autonomous driving scenarios and cannot meet general application requirements.
By adjusting the virtual camera's settings and determining the image information, an initial simulated video stream is generated, and then converted using the target video generation model and target style image to generate a target simulated video stream so that its image style is consistent with the real-world image.
The generated target simulation video stream is more realistic in picture content and style, and can be applied to a wide range of application scenarios, enriching the application scenarios of the digital simulation world, reducing costs and improving picture realism.
Smart Images

Figure IB2024062687_09102025_PF_FP_ABST
Abstract
Description
[0001]TECHNICAL FIELD The present disclosure relates to the fields of large-scale model technology and video processing technology, and more specifically, to a video stream generation method, electronic device, and computer-readable storage medium. Background: A digital twin system is a virtual simulation system based on a digital model that combines real-world entities or systems with the digital model to construct a virtual digital twin world corresponding to the real world. The digital twin system can use a rendering engine to generate virtual video streams that simulate the real world, thereby simulating and analyzing the real world. Currently, the CARLA simulator has been proposed for the autonomous driving field. Leveraging the powerful rendering engine, the CARLA simulator can generate virtual video streams that simulate real scenes. However, the CARLA simulator is only targeted at autonomous driving scenarios and is not suitable for general applications, resulting in a very limited range of applications. Furthermore, the CARLA simulator relies on the capabilities of the rendering engine, and the realism of the virtual video streams it generates needs to be improved. Currently, no effective solution has been proposed to address the aforementioned issues. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a video stream generation method, an electronic device, and a computer-readable storage medium to at least address the technical issue of low fidelity in virtual video streams. According to one aspect of the embodiments of the present disclosure, a video stream generation method is provided, comprising: obtaining setting information of a virtual camera, wherein the virtual camera is used to record a virtual video stream in a digital twin world (a virtual world that simulates the real world), and the setting information is used to adjust the shooting angle of the virtual camera; determining image information in a target simulated video stream to be generated, wherein the image information is used to record simulated image elements to be produced in the target simulated video stream to be generated; generating an initial simulated video stream based on the setting information and the image information; and converting the initial simulated video stream based on a target video generation model and a target style image to generate a target simulated video stream, wherein the target style image is a real image in the real world, and the style of the target simulated video stream corresponds to the style of the target style image.According to another aspect of an embodiment of the present disclosure, a video stream generation method is provided, comprising: in response to an adjustment instruction applied on an operation interface, displaying setting information of a virtual camera on the operation interface, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world simulating the real world, and the setting information is used to adjust the shooting angle of the virtual camera; in response to an input instruction applied on the operation interface, displaying picture information of a target simulated video stream to be generated on the operation interface, wherein the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated; and in response to a simulation instruction applied on the operation interface, displaying the target simulated video stream on the operation interface, wherein the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world, the picture style of the target simulated video stream corresponds to the picture style of the target style image, and the initial simulated video stream is generated based on the setting information and the picture information. According to another aspect of an embodiment of the present disclosure, a video stream generation method is provided, comprising: obtaining a simulated video stream generation request through a first application programming interface; and returning a simulated video stream generation response through a second application programming interface. The simulated video stream generation request includes: setting information of a virtual camera and image information of a target simulated video stream to be generated; and the simulated video stream generation response includes: a target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, where the digital twin world is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; the image information is used to record simulated image elements that need to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image, where the target style image is a real image in the real world; the image style of the target simulated video stream corresponds to the image style of the target style image; and the initial simulated video stream is generated based on the setting information and the image information.According to another aspect of an embodiment of the present disclosure, a video stream generation method is provided, comprising: obtaining an input simulated video stream generation dialog request; returning a simulated video stream generation dialog reply in response to the simulated video stream generation dialog request; wherein the simulated video stream generation dialog request includes: setting information of a virtual camera and image information of a target simulated video stream to be generated; and the simulated video stream generation dialog reply includes: a target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, where the digital twin world is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; and the image information is used to record simulated image elements that need to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image, where the target style image is a real image in the real world, the image style of the target simulated video stream corresponds to the image style of the target style image, and the initial simulated video stream is generated based on the setting information and the image information; and displaying the target simulated video stream in a graphical user interface. According to another aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a memory storing an executable program; and a processor configured to execute the program, wherein when the program is executed, any one of the aforementioned methods for generating a video stream is executed. According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores the executable program, wherein when the executable program is executed, the computer-readable storage medium controls the device containing the computer-readable storage medium to execute any one of the aforementioned methods for generating a video stream. According to another aspect of the embodiments of the present disclosure, a computer program product is provided, comprising the computer program, wherein when executed by the processor, the computer program implements any one of the aforementioned methods for generating a video stream.In the embodiments of the present disclosure, by adjusting the virtual camera's setting information, the shooting angle of the virtual camera recording the virtual video stream in the digital twin world is adjusted, and the image information required to record the target video stream to be generated is determined. Based on the setting information and image information, an initial simulated video stream is generated. The target style image in the real world and the initial simulated video stream are input into a target video generation model. Based on the target video generation model and the target style image, the initial simulated video stream is converted to generate the target simulated video stream. The style of the target simulated video stream corresponds to the style of the target style image. This achieves the goal of generating a target simulated video stream with specified content and a specified style. This achieves the goal of generating a target simulated video stream with more realistic and reliable content, a more realistic style, and applicability to any scenario. This also enriches the technical effects of the application scenarios in the digital twin world and solves the technical problem of low image fidelity in virtual video streams. It should be noted that the general description above and the detailed description that follow are merely illustrative and illustrative of the present disclosure and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure. In the accompanying drawings: Figure 1 is a schematic diagram of an application scenario of a video stream generation method according to Example 1 of the present disclosure; Figure 2 is a flow chart of a video stream generation method according to Example 1 of the present disclosure; Figure 3 is a flow chart of a training encoder according to Example 1 of the present disclosure; Figure 4 is a flow chart of a training decoder according to Example 1 of the present disclosure; Figure 5 is a flow chart of the inference process of the target video generation model according to Example 1 of the present disclosure; Figure 6 is a flow chart of another video stream generation method according to Example 1 of the present disclosure; Figure 7 is a flow chart of a video stream generation method according to Example 2 of the present disclosure; Figure 8 is a flow chart of a video stream generation method according to Example 3 of the present disclosure; Figure 9 is a flow chart of a video stream generation method according to Example 4 of the present disclosure; Figure 10 is a structural schematic diagram of a video stream generation device according to Example 5 of the present disclosure; Figure 11 is a structural schematic diagram of another video stream generation device according to Example 5 of the present disclosure; Figure 12 is a structural schematic diagram of another video stream generation device according to Example 5 of the present disclosure; Figure 13 is a structural schematic diagram of yet another video stream generation device according to Example 5 of the present disclosure; Figure 14 is a structural block diagram of an electronic device according to an embodiment of the present disclosure.DETAILED DESCRIPTION To help those skilled in the art better understand the present disclosure, the following will provide a clear and complete description of the technical solutions in the embodiments of the present disclosure, in conjunction with the accompanying drawings. It should be noted that the described embodiments represent only a portion of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort should fall within the scope of protection of the present disclosure. It should be noted that the terms "first," "second," and so on, in the specification and claims of the present disclosure, and in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the steps or units expressly listed, but may include other steps or units not expressly listed or inherent to such process, method, product, or apparatus. The technical solutions provided herein are primarily implemented using large-scale model technology. A large-scale model here refers to a deep learning model with large-scale model parameters, typically including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large-scale models, also known as foundation models, are pre-trained using large-scale unlabeled corpora to produce pre-trained models with parameters exceeding 100 million. Such models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities, such as large language models (LLMs) and multi-modal pre-training models. It should be noted that in practical applications, large-scale models can be fine-tuned using a small number of samples, allowing them to be applied to different tasks.For example, large models can be widely applied in fields such as natural language processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation. They can also be widely applied to natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios of large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. First, some nouns or terms used in the description of the embodiments of this disclosure are subject to the following interpretation: Digital twin system: refers to the use of digital technology to simulate and emulate real-world physical systems to achieve a digital mirroring of the actual physical system. The digital twin system uses sensors and data acquisition technology to convert the state and behavior of the real-world physical system into digital data in real time. Then, using computer simulation and emulation technology, this data is reconstructed into a digital twin model to accurately describe and predict the physical system. A digital twin world is a virtual environment constructed by a digital twin system. Using digital twin technology, real-world physical systems are digitized and simulated in a virtual environment, enabling real-time monitoring, prediction, and control of the real world. Digital twin worlds can help people better understand and analyze complex real-world systems, improving the accuracy and efficiency of decision-making. They also provide powerful tools and support for scientific research, engineering design, and manufacturing. The whitening and coloring transform (WCT) method is an image style transfer algorithm used to transfer the style of one image to another. This algorithm first whitens the input image to remove color and contrast information, then recolors it according to the style of the target image, achieving style transfer. Real video streams are generated by optically imaging the real world with a real camera. Although real video streams can be generated continuously, they are representations of the real world and cannot meet the specific requirements of digital twin worlds for virtual video streams.For example, people hope to simulate various accidents in a digital twin world to explore possible solutions. However, real-world video streams that accurately record accidents are extremely rare, and real video streams cannot record events that only exist in people's imaginations but have never occurred. Furthermore, real video streams require manual labeling, which consumes a significant amount of manpower and time, significantly limiting the ability to use massive amounts of labeled video stream data for model training. Furthermore, the label distribution of real video streams often exhibits errors, which are also detrimental to model training. Currently, the CARLA simulator has been proposed for the autonomous driving field. By leveraging the capabilities of a powerful rendering engine, the CARLA simulator can generate virtual video streams that simulate real autonomous driving scenarios. However, the CARLA simulator is designed specifically for autonomous driving scenarios and is not suitable for general applications, resulting in very limited application scenarios. Furthermore, the CARLA simulator relies on the capabilities of the rendering engine, and the realism of the virtual video streams it generates needs to be improved. Relying on the rendering engine to generate virtual video streams has the following drawbacks. Defect 1: Limited application scenarios. The currently proposed CARLA simulator is only applicable to autonomous driving scenarios, not general application scenarios, and does not meet the requirements of digital twin tasks. Defect 2: The virtual video stream generated by relying on the rendering engine's capabilities lacks realism. Prior to the present disclosure, no effective solution to these drawbacks had been proposed. Example 1: According to an embodiment of the present disclosure, a method for generating a video stream is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions. Furthermore, although the flowcharts illustrate a logical order, in some cases, the steps shown or described may be executed in a different order. Considering the large number of model parameters in large models and the limited computing resources of mobile terminals, the video stream generation method provided in the embodiments of the present disclosure can be applied to the application scenario shown in Figure 1, but is not limited thereto. In the application scenario shown in FIG1 , a large model is deployed on a server 10. Server 10 can be connected to one or more client devices 20 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Client devices 20 herein include, but are not limited to, smartphones, tablet computers, laptop computers, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users via a graphical user interface to invoke the large model and thereby implement the methods provided in the embodiments of the present disclosure.In an embodiment of the present disclosure, a system consisting of a client device and a server can perform the following steps: the client device sends a simulated video stream generation request for generating a target simulated video stream; the server adjusts virtual camera settings, determines image information in the target simulated video stream to be generated, generates an initial simulated video stream based on the settings and image information, and converts the initial simulated video stream based on a target video generation model and a target style image to generate the target simulated video stream. It should be noted that if the client device's operating resources meet the deployment and operating conditions of a large model, the present disclosure can be performed on the client device. In this operating environment, the present disclosure provides a video stream generation method as shown in Figure 2. Figure 2 is a flow chart of a video stream generation method according to Example 1 of the present disclosure. As shown in Figure 2, the method may include the following steps: Step S21: Acquiring setting information for a virtual camera, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setting information is used to adjust the shooting angle of the virtual camera; Step S22: Determining the image information in a target simulated video stream to be generated, wherein the image information is used to record the simulated image elements that need to be produced in the target simulated video stream to be generated; Step S23: Generating an initial simulated video stream based on the setting information and the image information. The setting information in this step may be dynamically adjusted setting information, i.e., the setting information is dynamically adjusted in the virtual camera to obtain the adjusted setting information. Step S24: Converting the initial simulated video stream based on a target video generation model and a target style image to generate a target simulated video stream, wherein the target style image is a real image in the real world, and the image style of the target simulated video stream corresponds to the image style of the target style image. In the embodiments of the present disclosure, a digital twin world refers to a virtual world corresponding to the real world, created through digitalization and simulation technologies. The digital twin world simulates the real world, enabling real-time monitoring, analysis, and prediction of the real world based on the digital twin world. A virtual camera is an important tool in the digital twin world, a virtual device that can simulate the functions of a real camera in the real world. The virtual camera is used to record a virtual video stream in the digital twin world. The virtual camera can capture scenes and data in the digital twin world, thereby recording a virtual video stream. The virtual video stream can be understood as virtual video data composed of simulated images. Based on the virtual video stream, various situations in the digital twin world can be more intuitively understood, providing important data support for monitoring, analysis, and prediction of the real world.The setting information can be understood as the setting parameters of the virtual camera, used to adjust the shooting angle of the virtual camera, that is, to change the shooting angle of the virtual video stream recorded by the virtual camera. Exemplarily, the setting information includes at least one of the following parameters: the virtual camera's position, posture, field of view, focal length, aperture, exposure time, and other parameters, which are not limited here. The target simulated video stream can be understood as the simulated video stream ultimately generated by the video stream generation method provided by the present disclosure. The screen information is used to record the simulated screen elements to be produced in the target simulated video stream to be generated. Simulated screen elements can be understood as the simulated elements that the user expects to appear in the ultimately generated target simulated video stream. Exemplarily, simulated screen elements can include, but are not limited to, objects in the virtual video stream, the paths of object elements, and the types of events described by the virtual video stream, which are not limited here. Object elements in the virtual video stream can be pedestrians, vehicles, animals, and the like. The paths of object elements can be pedestrian paths, vehicle paths, animal paths, and the like. The event types described by the virtual video stream can be common real-world events such as pedestrians crossing the road, vehicles turning left, and animals running. They can also be real-world event types with an extremely low probability of occurrence, or event types that have never occurred in the real world, i.e., event types that cannot be perceived or failed to be perceived in the real world, without limitation. By determining the image information in the target simulated video stream to be generated, the simulated image content to be included in the target simulated video stream can be determined. It is understood that the image information can be personalized according to the user's actual needs and can be applied to any scenario, without limitation, so that the final generated target simulated video stream can contain the simulated image content desired by the user. An initial simulated video stream can be generated based on the virtual camera setting information and the determined image information. For example, a basic rendering engine can be used to generate the initial simulated video stream based on the setting information and image information, eliminating the need for a powerful rendering engine and thus saving costs. The initial simulated video stream can be understood as a basic virtual video stream used to simulate a real-world video stream. Its image quality is relatively coarse, its style is monotonous, and its visual fidelity is low. Therefore, the initial simulated video stream requires further processing. The target style image is a real image in the real world, which can be understood as a real image that can be captured by a real camera in the real world. The target style image can be personalized based on the user's actual needs and is not limited here. The target video generation model can be a large model or other deep learning model and is not limited here.The target video generation model transforms the initial simulated video based on the input target style image. Specifically, it transforms the initial simulated video's visual style, generating a target simulated video stream that matches the style of the target style image. This style is then matched to the target style image, resulting in a target simulated video stream. This ensures that the target simulated video stream's visual style matches that of the target style image, ensuring a consistent visual style. This results in a more realistic target simulated video stream and a better simulation of the real video stream. Furthermore, this model avoids the inclusion of content in the target simulated video stream that does not exist in the real world. In the disclosed embodiments, the virtual camera's setting information is adjusted to adjust the shooting angle of the virtual camera recording the virtual video stream in the digital twin world. The image information required to record the target simulated video stream to be generated is determined. Based on the setting information and the image information, an initial simulated video stream is generated. The target style image in the real world and the initial simulated video stream are then input into a target video generation model. The initial simulated video stream is then converted based on the target video generation model and the target style image, ultimately generating the target simulated video stream. This ensures that the image style of the target simulated video stream corresponds to that of the target style image. As can be seen, the disclosed embodiments enable personalized settings for the simulated image content in the target simulated video stream, ensuring that the target simulated video stream includes the simulated image elements recorded by the image information. This allows for a wide range of applications. Furthermore, the image style of the target simulated video stream can be made consistent with that of the target style image in the real world, resulting in a higher degree of realism and a closer approximation to reality, thereby allowing users to experience both authenticity and diversity. Furthermore, the present disclosure can generate a target simulated video stream with specified content and realistic style in any scenario in the digital twin world. This can simulate a perceptual video stream that meets the requirements of the digital twin task, greatly enriching the application scenarios of the digital twin task. The video stream generation method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving simulated video stream generation in fields such as e-commerce services, educational services, legal services, medical services, conference services, social networking services, financial product services, logistics services, and navigation services. For example, simulated video stream generation for e-commerce services, educational services, and legal services, etc., are not limited here.According to the disclosed embodiments, by identifying the setting information of a virtual camera, the shooting angle of the virtual camera used to record a virtual video stream in a digital twin world is adjusted, and the image information required to be produced in the target simulated video stream to be generated is determined. Based on the setting information and the image information, an initial simulated video stream is generated. A target style image in the real world and the initial simulated video stream are input into a target video generation model. The initial simulated video stream is converted based on the target video generation model and the target style image to ultimately generate the target simulated video stream. The style of the target simulated video stream corresponds to the style of the target style image, thereby achieving the goal of generating a target simulated video stream with specified image content and a specified style. This achieves the goal of generating a target simulated video stream with more realistic and reliable image content, a more realistic image style, and applicability to any scenario. This also enriches the technical effects of the application scenarios in the digital twin world and solves the technical problem of low image fidelity in the virtual video stream. In an optional embodiment, in step S23, an initial simulated video stream is generated based on the setting information and the image information, including the following method steps: generating a target tag based on the setting information and the image information, wherein the target tag is used to identify the video attributes of the target simulated video stream to be generated; rendering the setting information and the image information using a digital twin system to obtain a first simulated video stream, wherein the digital twin system is used to construct a digital twin world; and labeling the first simulated video stream with the target tag to obtain an initial simulated video stream. In the disclosed embodiment, the target tag is used to identify the video attributes of the target simulated video stream to be generated. The video attributes of the target simulated video stream may include, but are not limited to, the target simulated video stream's screen size, screen orientation, screen style, screen content, and the event type described by the target simulated video stream, and are not limited here. The digital twin system can be understood as a digital system for constructing a digital twin world. A basic rendering engine is used in the digital twin system to render the virtual camera setting information and the determined screen information to obtain the first simulated video stream. It is understandable that the image quality of the first simulated video stream is relatively rough and the screen style is monotonous. In an embodiment of the present disclosure, when generating an initial simulated video stream based on setting information and picture information, a target tag for identifying the video attributes of the target simulated video stream to be generated can be generated based on the setting information and picture information. Simultaneously, the setting information and picture information are rendered using a digital twin system to obtain a first simulated video stream. The first simulated video stream is then labeled with the target tag to obtain the initial simulated video stream, thereby achieving the technical effect of automatically labeling the initial simulated video stream with the target tag.It is understandable that, in one embodiment, video streams typically require manual labeling, resulting in high labor and time costs. In the disclosed embodiment, a target label for identifying the video attributes of a target simulated video stream to be generated can be automatically generated based on the setting information and the image information. The target label is then used to label the first simulated video stream, thereby generating an initial simulated video stream. This initial simulated video stream carries the target label, eliminating the need for manual labeling of the video stream, effectively reducing labor and time costs. In an optional embodiment, the video stream generation method further includes the following method steps: obtaining a first image sample and a second image sample, wherein the first image sample is a real image sample in the real world, and the second image sample is a virtual image sample corresponding to the image content of the first image sample; and training an initial video generation model based on the first and second image samples to obtain a target video generation model. In the disclosed embodiment, the first image sample can be understood as a real image sample captured by a real camera in the real world, and the second image sample can be understood as a virtual image sample in the virtual world, corresponding to the image content of the first image sample. For example, the first image sample may be a real puppy image sample, and the second image sample may be a virtual puppy image sample corresponding to the image content of the first image sample. It should be noted that the sample magnitudes of the real and virtual image samples described above are sufficient for training a video generation model, and this is not a limitation. The initial video generation model is the model of the target video generation model before training, and may be a large model or other deep learning model, and this is not a limitation. In embodiments of the present disclosure, when training the target video generation model, a first image sample captured in the real world and a second image sample corresponding to the image content of the first image sample are obtained. The initial video generation model is then trained based on the first and second image samples to obtain the target video generation model. In an optional embodiment, obtaining the second image sample includes the following method steps: performing image detection on the first image sample to obtain a first detection result, wherein the first detection result represents real image elements in the first image sample; and performing image simulation processing on the first detection result to generate a second image sample, wherein the simulated image elements in the second image sample correspond to the real image elements in the first image sample.In the embodiments of the present disclosure, performing image detection on the first image sample can be understood as performing image content detection on the first image sample to detect the real screen elements included in the first image sample, thereby obtaining a first detection result. The first detection result is used to represent the real screen elements in the first image sample. Performing image simulation processing on the first detection result can be understood as performing image simulation processing on the real screen elements represented by the first detection result to simulate and obtain simulated screen elements corresponding to the real screen elements, and then generating a second image sample including the simulated screen elements. The simulated screen elements in the second image sample correspond to, i.e., remain consistent with, the real screen elements in the first image sample. In the embodiments of the present disclosure, when obtaining the second image sample, image detection can be performed on the first image sample to obtain a first detection result representing the real screen elements in the first image sample. Image simulation processing is then performed on the first detection result to generate the second image sample. The simulated screen elements in the second image sample correspond to the real screen elements in the first image sample, i.e., the simulated screen elements in the second image sample are kept consistent with the real screen elements in the first image sample as much as possible. In an optional embodiment, obtaining a second image sample includes the following method steps: performing image detection on the first image sample to obtain a second detection result, wherein the second detection result is used to represent the content of the real event described by the first image sample; and performing image simulation processing on the second detection result to generate a second image sample, wherein the simulated event content described by the second image sample corresponds to the content of the real event described by the first image sample. In the embodiment of the present disclosure, performing image detection on the first image sample can be understood as performing event content detection on the first image sample to obtain the content of the real event described by the first image sample, thereby obtaining the second detection result, which is used to represent the content of the real event described by the first image sample. Performing image simulation processing on the second detection result can be understood as performing image simulation processing on the real event content represented by the second detection result to obtain simulated event content corresponding to the real event content, thereby generating a second image sample describing the simulated event content. The simulated event content described by the second image sample corresponds to, i.e., remains consistent with, the content of the real event described by the first image sample.In an embodiment of the present disclosure, when obtaining a second image sample, image detection can be performed on the first image sample to obtain a second detection result representing the content of the real event described by the first image sample. Image simulation processing is then performed on the second detection result to generate a second image sample, such that the content of the simulated event described by the second image sample corresponds to the content of the real event described by the first image sample, that is, the content of the simulated event described by the second image sample is kept consistent with the content of the real event described by the first image sample as much as possible. In an optional embodiment, an initial video generation model is trained based on the first and second image samples to obtain a target video generation model. The method includes the following steps: performing image perspective conversion on the first image sample to generate a third image sample, where the third image sample is an image of the first image sample at a different perspective; performing image perspective conversion on the second image sample to generate a fourth image sample, where the fourth image sample is an image of the second image sample at a different perspective; and training the initial video generation model based on the first, second, third, and fourth image samples to obtain a target video generation model. In the disclosed embodiment, after obtaining the first and second image samples, the first image sample may be subjected to image perspective conversion to generate a third image sample. This third image sample is an image of the first image sample at a different perspective. Simultaneously, the second image sample may be subjected to image perspective conversion to generate a fourth image sample. This fourth image sample is an image of the second image sample at a different perspective. This allows the target video generation model to be trained based on real image samples at multiple perspectives and corresponding virtual image samples at multiple perspectives, thereby improving the output accuracy of the trained target video generation model. For example, when performing image perspective conversion on the first image sample, the first image sample may be subjected to operations such as rotation, subtraction, and Gaussian noise addition to generate the third image sample at a different perspective. This is not intended to be limiting. In an embodiment of the present disclosure, when an initial video generation model is trained based on a first image sample and a second image sample to obtain a target video generation model, an image perspective conversion may be performed on the first image sample to generate a third image sample of the first image sample at a different perspective, and an image perspective conversion may be performed on the second image sample to generate a fourth image sample of the second image sample at a different perspective. Then, the initial video generation model is trained based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain the target video generation model.In an optional embodiment, training an initial video generation model based on first, second, third, and fourth image samples to obtain a target video generation model includes the following method steps: training an initial encoder based on the first, second, third, and fourth image samples to obtain a target encoder; training an initial decoder based on the first and third image samples to obtain a target decoder; and determining a target video generation model based on the target encoder and target decoder. In the disclosed embodiment, when training the initial video generation model based on the first and second image samples to obtain the target video generation model, the initial encoder may be trained based on the first, second, third, and fourth image samples, that is, the initial encoder is trained based on real image samples from different perspectives and corresponding virtual image samples, thereby obtaining the target encoder. Simultaneously, the initial decoder is trained based on the first and third image samples, that is, the initial decoder is trained based on real image samples from different perspectives as input and output supervision signals, thereby obtaining the target decoder. Finally, a target video generation model is determined based on the target encoder and target decoder. Specifically, the parameters of the trained target encoder and target decoder are frozen, and the target encoder and target decoder are connected so that the output of the target encoder serves as the input of the target decoder, thereby forming a complete model, namely, the target video generation model. In an optional embodiment, an initial encoder is trained based on first, second, third, and fourth image samples to obtain a target encoder. The method includes the following steps: determining the first and third image samples as a first sample combination; determining the second and fourth image samples as a second sample combination; and determining the first and second image samples as a third sample combination; and performing contrastive learning training on the initial encoder based on a contrastive loss function and the first, second, and third sample combinations to obtain the target encoder. In the disclosed embodiment, when training the initial encoder based on the first, second, third, and fourth image samples to obtain the target encoder, real image samples from different perspectives can be paired and combined with virtual image samples to obtain the first, second, and third sample combinations. Then, based on the contrastive loss function, the first sample combination, the second sample combination, and the third sample combination, contrastive learning training is performed on the initial encoder to obtain the target encoder.A first sample combination is determined based on the first and third image samples. The first sample combination is a combination of real image samples from different perspectives. A second sample combination is determined based on the second and fourth image samples. The second sample combination is a combination of virtual image samples from different perspectives. A third sample combination is determined based on the first and second image samples. The third sample combination is a combination of real image samples and corresponding virtual image samples from the same perspective. For example, taking the example where the first sample combination includes an image sample from a first perspective and an image sample from a second perspective, when contrastive learning is performed on the initial encoder based on the contrastive loss function, the first sample combination, the second sample combination, and the third sample combination, the initial encoder can be used to encode the image sample from the first perspective to obtain a first feature, and to encode the image sample from the second perspective to obtain a second feature. Because the image samples from the first and second perspectives have similar image content, the contrastive learning loss function is used to shorten the feature distance between the first and second features and to increase the distance between the first feature and features of other samples, thereby performing contrastive learning training on the initial encoder. Similarly, a similar process is performed on the second and third sample combinations (not described in detail here), thereby bringing the feature distances between real image samples with similar image content and the corresponding virtual image samples closer, thereby training the target encoder. Figure 3 is a flowchart of encoder training according to Embodiment 1 of the present disclosure. As shown in Figure 3, during encoder training, a mixed set of real image samples and virtual image samples is divided according to view angle to obtain view angle 1 and view angle 2. In this embodiment of the present disclosure, there are three combinations of pairings of view angle 1 and view angle 2: the first, second, and third sample combinations described above. The encoder is then trained using contrastive learning based on the contrastive loss function and view angles 1 and 2. In an optional embodiment, an initial decoder is trained based on the first and third image samples to obtain a target decoder, including the following method steps: The initial decoder is trained autoregressively based on the first and third image samples to obtain the target decoder. In the disclosed embodiment, when the initial decoder is trained based on the first and third image samples to obtain the target decoder, the initial decoder can be trained autoregressively based on the first and third image samples to obtain the target decoder. That is, real image samples from multiple perspectives are used as input and output supervisory signals to perform autoregressive training on the initial decoder, thereby obtaining the target decoder.FIG4 is a flowchart of training a decoder according to Example 1 of the present disclosure. As shown in FIG4 , the parameters of a trained encoder can be frozen. A real image sample is then input into the encoder for feature encoding. The resulting encoding result is then input into the decoder, which decodes the encoding result and outputs a target image. The output target image is then compared with the input real image sample, the difference between the two is calculated, and the decoder parameters are adjusted based on the difference. Autoregressive training is then performed on the decoder until the decoder can output a target image that meets the requirements. In an optional embodiment, in step S24, an initial simulated video stream is converted based on a target video generation model and a target style image to generate a target simulated video stream. The method includes the following steps: using a target encoder to perform feature encoding on the target style image to obtain a first encoding result, and using the target encoder to perform feature encoding on video frames of the initial simulated video stream to obtain a second encoding result; performing feature fusion on the first encoding result and the second encoding result to obtain a fused result; decoding the fused result using a target decoder to obtain the target image; and generating a target simulated video stream based on the target image. In an embodiment of the present disclosure, when converting an initial simulated video stream based on a target video generation model and a target style image to generate a target simulated video stream, the target style image and the initial simulated video stream can be input into the target video generation model. A target encoder in the target video generation model then performs feature encoding on the target style image to obtain a first encoding result. The target encoder then performs feature encoding on video frames of the initial simulated video stream to obtain a second encoding result. Specifically, the initial simulated video stream is processed frame by frame, and feature encoding is performed on each video frame to obtain the second encoding result. After obtaining the first encoding result and the second encoding result, the first encoding result and the second encoding result are feature fused to obtain a fused result, i.e., a new feature encoding having the simulated content of the initial simulated video stream and the style of the target style image. The fused result is then decoded using a target decoder in the target video generation model to obtain a target image. Finally, a target simulated video stream is generated based on the target image, such that the target simulated video stream has specified simulated content and a specified style. Figure 5 is a flow chart of the inference process of the target video generation model according to Example 1 of the present disclosure. As shown in Figure 5, the virtual video frame image of the initial simulated video stream input into the target video generation model and the real target style image are first feature-encoded by the target encoder to obtain the encoding results of the two.The encoding results of the two are then subjected to feature fusion based on a fusion module. For example, a whitening and coloring transform (WCT) method can be used to fuse the features of the two encoding results to obtain a fused result. The fused result is then decoded by a target decoder, thereby outputting a target image having the content of the virtual video frame image and the style of the target style image. Finally, a target simulated video stream is constructed based on the output target image. In an optional embodiment, in step S21, obtaining virtual camera setting information includes the following method steps: determining the virtual camera's position information, posture information, and field of view angle information; and obtaining the setting information based on the position information, posture information, and field of view angle information. In the disclosed embodiment, the virtual camera's position information can be understood as the position coordinates of the virtual camera in the digital twin world, denoted as (x, y, z), where x can be understood as the longitude of the virtual camera, y can be understood as the latitude of the virtual camera, and z can be understood as the altitude of the virtual camera. The virtual camera's pose information can be understood as the virtual camera's orientation in the digital twin world, denoted as (h, p, r), where h represents the virtual camera's yaw angle, p represents its pitch angle, and r represents its roll angle. The virtual camera's field of view (FOV) information can be understood as the range of the virtual camera's observation, and can be expressed as an angle. It is understood that a larger FOV indicates a wider range of observation, while a smaller FOV indicates a narrower range. In the disclosed embodiments, when acquiring virtual camera configuration information, the virtual camera's position information, pose information, and FOV information can be determined based on actual needs. The virtual camera's configuration information is then acquired based on this position information, pose information, and FOV information, ensuring that the virtual camera's shooting angle meets actual needs.In an optional embodiment, the image information includes the target event type and event duration of the target event described by the target simulation video stream to be generated. The target event type includes an event type that indicates a perception failure in the real world. The event duration represents the duration from the onset to the end of the target event. Using a digital twin system to render the setting information and image information to obtain a first simulation video stream includes the following method steps: First, determining the video duration of the first simulation video based on the event duration, wherein the video duration represents the duration from the start to the end of the first simulation video playback, and the video duration is greater than or equal to the event duration; Then, using the digital twin system to render the video duration, setting information, and image information to obtain an initial simulation video stream. In the disclosed embodiment, the image information may include, but is not limited to, the target event type and event duration of the target event described by the target simulation video stream to be generated, object elements in the target simulation video stream, and the motion paths of the object elements, etc., without limitation herein. The event duration of the target event can be understood as the duration from the onset to the end of the target event. Target event types may include, but are not limited to, event types that cannot be perceived (i.e., perception failures) in the real world, event types that frequently occur in the real world, and event types with an extremely low probability of occurring in the real world, without limitation here. In an embodiment of the present disclosure, when using the digital twin system to render the setting information and image information to obtain a first simulated video stream, the video duration of the first simulated video can be first determined based on the event duration of the target event recorded in the image information. This video duration must be greater than or equal to the event duration. The digital twin system is then used to render the video duration, setting information, and image information to obtain an initial simulated video stream. Figure 6 is a flow chart of another video stream generation method according to Example 1 of the present disclosure. As shown in Figure 6, the camera position, posture, field of view, and other information of the virtual camera in the digital twin world are first determined and set in the digital twin system. The image content to appear in the generated simulated video stream is then set according to user requirements, such as pedestrians, vehicle trajectories, and event types. Simulated labels can also be generated based on the set image content and the determined virtual camera parameters. Next, a base video stream is recorded using the digital twin engine, based on the duration of the target event set in the image content. Finally, the recorded base video stream and the specified style image are input into the target video generation model, transforming the base video stream into a realistic video stream with the user-specified realistic style. As can be seen, the disclosed embodiments have developed and implemented a video stream perception simulation system based on the digital twin world.This system utilizes a basic rendering engine to generate a basic virtual video stream and further leverages image generation technology to convert this basic virtual video stream into a realistic video stream with a specified style. The content of this generated video stream can be customized to simulate various scenarios based on task requirements and comes with its own attribute annotations. This video stream perception simulation system can simulate perceptual video streams that meet the functional requirements of digital twins, significantly enriching the application scenarios of digital twins and reducing labor and time costs. Compared to one implementation that relies on a powerful rendering engine but only generates virtual video streams with a single visual style and low realism, the present disclosure, based on a basic rendering engine and a target video generation model, enables targeted generation according to user-specified style types, resulting in a more diverse and realistic target simulated video stream. Furthermore, while one embodiment utilizes the renderer's capabilities to alter the unrealistic content of a virtual video stream, the present disclosure can generate a target simulated video stream with specified content, making the content more realistic and reliable. Furthermore, it can generate simulated video streams corresponding not only to commonly occurring real-world events but also to highly unlikely or previously unoccurred real-world events, significantly enriching the scenarios of the digital twin world. Furthermore, while one embodiment requires manual labeling of video streams, the target simulated video stream generated by the present disclosure comes with built-in target labels, which can be used to identify the video attributes of the target simulated video stream, thereby reducing labor costs and time, and enabling model training based on massive amounts of labeled video data. It is easy to understand that the video stream generation method provided by the present disclosure has the following beneficial effects. The present disclosure only requires the use of a basic renderer to generate a base video stream. This base video stream is then targeted and transformed using a target simulated video generation model and a specified target image style, resulting in a more realistic target simulated video stream with more realistic and reliable content. The target simulation video streams generated by this disclosure can be customized to simulate various scenarios based on task requirements, making them suitable for a wide range of applications. This technology can generate target simulation video streams corresponding not only to commonly occurring real-world events but also to highly unlikely or previously unoccurred real-world events, significantly enriching the scenarios of the digital twin world. The target simulation video streams generated by this disclosure come with built-in target labels, which can be used to identify their video attributes. This reduces labor and time costs and makes it possible to train models based on massive amounts of labeled video data.It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. Furthermore, it should be noted that for the sake of simplicity, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited by the order of the actions described, as certain steps can be performed in a different order or simultaneously according to this disclosure. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required by this disclosure. Through the above description of the embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software and a necessary general-purpose hardware platform, or alternatively, hardware. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present disclosure.Example 2 In the operating environment of Example 1, the present disclosure provides a video stream generation method as shown in Figure 7. Figure 7 is a flowchart of a video stream generation method according to Example 2 of the present disclosure. As shown in Figure 7, the method includes: Step S71, in response to an adjustment instruction applied to an operation interface, displaying setting information of a virtual camera on the operation interface, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world, and the setting information is used to adjust the shooting angle of the virtual camera; Step S72, in response to an input instruction applied to the operation interface, displaying picture information in a target simulation video stream to be generated on the operation interface, wherein the picture information is used to record simulation picture elements that need to be produced in the target simulation video stream to be generated; Step S73, in response to a simulation instruction applied to the operation interface, displaying the target simulation video stream on the operation interface, wherein the target simulation video stream is generated by converting an initial simulation video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world, and the picture style of the target simulation video stream corresponds to the picture style of the target style image. The initial simulated video stream is generated based on the setting information and the screen information. In the embodiments of the present disclosure, an adjustment instruction can be understood as an instruction for adjusting the setting parameters of the virtual camera, triggered by a user performing an adjustment operation on the operation interface. An input instruction can be understood as an instruction for displaying the screen information of the target simulated video stream to be generated, triggered by a user performing an input operation on the operation interface. A simulation instruction can be understood as an instruction for displaying the target simulated video stream, triggered by a user performing a simulation operation on the operation interface. In the embodiments of the present disclosure, upon receiving an adjustment instruction on the operation interface, the setting information of the virtual camera can be displayed on the operation interface, thereby adjusting the setting parameters of the virtual camera. The virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setting information is used to adjust the shooting angle of the virtual camera. Upon receiving an input instruction on the operation interface, the screen information of the target simulated video stream to be generated can be displayed on the operation interface. The screen information records the simulated screen elements to be produced in the target simulated video stream to be generated.Upon receiving a simulation instruction applied to the operation interface, a target simulated video stream can be displayed on the operation interface. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image. The target style image is a real-world image, and the style of the target simulated video stream corresponds to that of the target style image. The initial simulated video stream is generated based on the setting information and the image information. As can be seen, the present disclosure can personalize the simulated image content in the target simulated video stream, ensuring that the target simulated video stream includes the simulated image elements recorded by the image information, thus embracing a wide range of application scenarios. The style of the target simulated video stream can also be aligned with that of the target style image in the real world, making the target simulated video stream more realistic and closer to reality, thereby allowing users to experience authenticity and diversity. Furthermore, the present disclosure can generate a target simulated video stream with specified content and a style close to reality in any scenario in the digital twin world. This allows for the simulation of perceptual video streams that meet the requirements of digital twin tasks, significantly enriching the application scenarios of digital twin tasks. The above-mentioned video stream generation method provided in the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving simulation video stream generation in the fields of e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services, for example: simulation video stream generation for e-commerce services, simulation video stream generation for educational services, simulation video stream generation for legal services, etc., which are not limited here.According to an embodiment of the present disclosure, upon receiving an adjustment instruction on an operation interface, setting information of a virtual camera can be displayed on the operation interface, thereby adjusting the setting parameters of the virtual camera. The virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setting information is used to adjust the shooting angle of the virtual camera. Upon receiving an input instruction on the operation interface, image information of a target simulated video stream to be generated can be displayed on the operation interface. The image information records the simulated image elements to be produced in the target simulated video stream to be generated. Upon receiving a simulation instruction on the operation interface, the target simulated video stream can be displayed on the operation interface. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image. The target style image is a real image in the real world. The image style of the target simulated video stream corresponds to the image style of the target style image. The initial simulated video stream is generated based on the setting information and the image information. This achieves the purpose of generating a target simulated video stream with specified image content and specified image style, thereby achieving more realistic and reliable image content. The target simulated video stream has a more realistic visual style and can be applied to any scenario. This enriches the technical effects of the application scenarios in the digital twin world, thereby resolving the technical issue of low visual fidelity in virtual video streams. It should be noted that the preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here. Example 3: In the operating environment of Example 1, the present disclosure provides a video stream generation method as shown in FIG8 .FIG8 is a flowchart of a video stream generation method according to Embodiment 3 of the present disclosure. As shown in FIG8 , the method includes: step S81, obtaining a simulated video stream generation request through a first application programming interface; step S82, returning a simulated video stream generation response through a second application programming interface; wherein the simulated video stream generation request includes: setting information of a virtual camera and image information of a target simulated video stream to be generated; and the simulated video stream generation response includes: a target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world; setting information is used to adjust the shooting angle of the virtual camera; image information is used to record simulated image elements that need to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image; the target style image is a real image in the real world; the image style of the target simulated video stream corresponds to the image style of the target style image; and the initial simulated video stream is generated based on the setting information and the image information. In the embodiment of the present disclosure, the simulation video stream generation request can be understood as a request for generating a simulation video stream, and the simulation video stream generation request includes: setting information of the virtual camera and picture information in the target simulation video stream to be generated. The simulation video stream generation response is a reply response corresponding to the simulation video stream generation request, and the simulation video stream generation response includes: the target simulation video stream. In the embodiment of the present disclosure, the simulation video stream generation request is obtained through the first application programming interface, wherein the simulation video stream generation request includes: setting information of the virtual camera and picture information in the target simulation video stream to be generated, the virtual camera is used to record the virtual video stream in the digital twin world, and the digital twin world is a virtual world that simulates the real world, the setting information is used to adjust the shooting angle of the virtual camera, and the picture information is used to record the simulation picture elements that need to be produced in the target simulation video stream to be generated. A simulated video stream generation response can then be returned via the second application programming interface. The simulated video stream generation response includes a target simulated video stream, generated by converting the initial simulated video stream based on a target video generation model and a target style image. The target style image is a real-world image. The style of the target simulated video stream corresponds to the style of the target style image. The initial simulated video stream is generated based on the setting information and the image information. As can be seen, the present disclosure can personalize the simulated image content in the target simulated video stream, ensuring that the target simulated video stream includes the simulated image elements specified by the image information, thus embracing a wide range of application scenarios.The target simulated video stream's visual style can also be made consistent with that of a target-style image in the real world, resulting in a higher degree of fidelity and closer to reality, allowing users to experience authenticity and diversity. Furthermore, the present disclosure can generate a target simulated video stream with specified visual content and a visual style close to reality in any scenario in the digital twin world. This can simulate a perceptual video stream that meets the requirements of digital twin tasks, greatly enriching the application scenarios of digital twin tasks. The video stream generation method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving simulated video stream generation in fields such as e-commerce services, educational services, legal services, medical services, conference services, social networking services, financial product services, logistics services, and navigation services. For example, simulated video stream generation for e-commerce services, educational services, and legal services, etc., are not limited here. In an embodiment of the present disclosure, a simulated video stream generation request is obtained through a first application programming interface (API). The simulated video stream generation request includes: virtual camera setup information and image information of a target simulated video stream to be generated. The virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setup information is used to adjust the shooting angle of the virtual camera. The image information is used to record the simulated image elements to be produced in the target simulated video stream to be generated. A simulated video stream generation response is then returned through a second application programming interface. The simulated video stream generation response includes: a target simulated video stream. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image. The target style image is a real image in the real world. The image style of the target simulated video stream corresponds to the image style of the target style image. The initial simulated video stream is generated based on the setup information and the image information. This achieves the purpose of generating a target simulated video stream with specified image content and specified image style, thereby achieving a target simulated video stream with more realistic and reliable image content, more realistic image style, and applicable to any scenario. At the same time, the technical effects of the application scenarios in the digital twin world are enriched, thereby solving the technical problem of low realism in the virtual video stream. It should be noted that the preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here. Example 4 In the operating environment of Example 1, the present disclosure provides a video stream generation method as shown in Figure 9.FIG9 is a flowchart of a video stream generation method according to Embodiment 4 of the present disclosure. As shown in FIG9 , the method includes: step S91, obtaining an input simulation video stream generation dialogue request; step S92, returning a simulation video stream generation dialogue reply in response to the simulation video stream generation dialogue request; the simulation video stream generation dialogue request includes: setting information of a virtual camera and picture information in a target simulation video stream to be generated; the simulation video stream generation dialogue reply includes: a target simulation video stream, where the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; the picture information is used to record the simulation picture elements that need to be produced in the target simulation video stream to be generated; the target simulation video stream is generated by converting an initial simulation video stream based on a target video generation model and a target style image, where the target style image is a real image in the real world, the picture style of the target simulation video stream corresponds to the picture style of the target style image, and the initial simulation video stream is generated based on the setting information and the picture information; step S93, The target simulated video stream is displayed within a graphical user interface. In the disclosed embodiments, a simulated video stream generation dialog request can be understood as a dialog request between a user and an intelligent machine. The simulated video stream generation dialog request includes virtual camera settings and image information of the target simulated video stream to be generated. A simulated video stream generation dialog response can be understood as a dialog response to the simulated video stream generation dialog request. The simulated video stream generation dialog response includes the target simulated video stream. In an embodiment of the present disclosure, a simulated video stream generation dialog request is generated by obtaining an input, wherein the simulated video stream generation dialog request includes: virtual camera setting information and image information of a target simulated video stream to be generated. Then, in response to the simulated video stream generation dialog request, a simulated video stream generation dialog reply is returned, wherein the simulated video stream generation dialog reply includes: the target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world; setting information is used to adjust the shooting angle of the virtual camera; image information is used to record the simulated image elements to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image; the target style image is a real image in the real world; the image style of the target simulated video stream corresponds to the image style of the target style image; and the initial simulated video stream is generated based on the setting information and the image information. Finally, the target simulated video stream is displayed in a graphical user interface to provide feedback to the user.As can be seen, the present disclosure can personalize the simulated image content in a target simulated video stream, ensuring that the target simulated video stream includes the simulated image elements recorded by the image information, thus enabling a wide range of application scenarios. It can also align the image style of the target simulated video stream with that of a target-style image in the real world, thereby enhancing the fidelity and approximation of the target simulated video stream, thereby allowing users to experience authenticity and diversity. Furthermore, the present disclosure can generate a target simulated video stream with specified image content and a style close to reality in any scenario in the digital twin world. This can simulate a perceptual video stream that meets the requirements of digital twin tasks, significantly enriching the application scenarios of digital twin tasks. The video stream generation method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving simulated video stream generation in fields such as e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services, for example, generating simulated video streams for e-commerce services, educational services, and legal services, without limitation here. According to an embodiment of the present disclosure, a simulated video stream generation dialog request is generated by obtaining an input simulated video stream, wherein the simulated video stream generation dialog request includes: virtual camera setting information and image information of a target simulated video stream to be generated. Then, in response to the simulated video stream generation dialog request, a simulated video stream generation dialog reply is returned, wherein the simulated video stream generation dialog reply includes: the target simulated video stream, wherein the virtual camera is used to record the virtual video stream in a digital twin world, which is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; the image information is used to record the simulated image elements to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting the initial simulated video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world; the image style of the target simulated video stream corresponds to the image style of the target style image, and the initial simulated video stream is generated based on the setting information and the image information. Finally, the target simulated video stream is displayed in a graphical user interface to provide feedback to the user, thereby achieving the purpose of generating a target simulated video stream with specified image content and specified image style. This enables the generation of a target simulated video stream with more realistic and reliable content, a more lifelike visual style, and applicability to any scenario. This also enriches the technical effects of the application scenarios in the digital twin world and resolves the technical issue of low visual fidelity in virtual video streams. It should be noted that the preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here.Example 5 According to an embodiment of the present disclosure, a device embodiment for implementing the above-mentioned video stream generation method is also provided. FIG10 is a schematic structural diagram of a video stream generating apparatus according to Embodiment 5 of the present disclosure. As shown in FIG10 , the video stream generating apparatus 1000 includes: a first acquiring module 1001, configured to acquire setting information of a virtual camera, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world simulating the real world, and the setting information is used to adjust the shooting angle of the virtual camera; a determining module 1002, configured to determine picture information in a target simulated video stream to be generated, wherein the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated; a first generating module 1003, configured to generate an initial simulated video stream based on the setting information and the picture information; and a second generating module 1004, configured to convert the initial simulated video stream based on a target video generation model and a target style image to generate a target simulated video stream, wherein the target style image is a real image in the real world, and the picture style of the target simulated video stream corresponds to the picture style of the target style image. Optionally, the first generation module 1003 is further configured to: generate a target label based on the setting information and the screen information, wherein the target label is used to identify the video attributes of the target simulated video stream to be generated; render the setting information and the screen information using a digital twin system to obtain a first simulated video stream, wherein the digital twin system is used to construct a digital twin world; and label the first simulated video stream with the target label to obtain an initial simulated video stream. Optionally, the apparatus further includes a training module configured to obtain a first image sample and a second image sample, wherein the first image sample is a real image sample in the real world and the second image sample is a virtual image sample corresponding to the screen content of the first image sample; and train the initial video generation model based on the first and second image samples to obtain a target video generation model. Optionally, the training module is further configured to: perform image detection on the first image sample to obtain a first detection result, wherein the first detection result is used to represent a real picture element in the first image sample; and perform image simulation processing on the first detection result to generate a second image sample, wherein the simulated picture elements in the second image sample correspond to the real picture elements in the first image sample.Optionally, the training module is further configured to: perform image detection on the first image sample to obtain a second detection result, wherein the second detection result is used to represent the content of the real event described by the first image sample; perform image simulation processing on the second detection result to generate a second image sample, wherein the content of the simulated event described by the second image sample corresponds to the content of the real event described by the first image sample. Optionally, the apparatus further includes: a conversion module configured to perform image perspective conversion on the first image sample to generate a third image sample, wherein the third image sample is an image of the first image sample at a different perspective; perform image perspective conversion on the second image sample to generate a fourth image sample, wherein the fourth image sample is an image of the second image sample at a different perspective; and train an initial video generation model based on the first, second, third, and fourth image samples to obtain a target video generation model. Optionally, the training module is further configured to: train an initial encoder based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain a target encoder; train an initial decoder based on the first image sample and the third image sample to obtain a target decoder; and determine a target video generation model based on the target encoder and the target decoder. Optionally, the training module is further configured to: determine the first image sample and the third image sample as a first sample combination; determine the second image sample and the fourth image sample as a second sample combination; and determine the first image sample and the second image sample as a third sample combination; perform contrastive learning training on the initial encoder based on the contrastive loss function, the first sample combination, the second sample combination, and the third sample combination to obtain the target encoder. Optionally, the training module is further configured to: perform autoregressive training on the initial decoder based on the first image sample and the third image sample to obtain the target decoder. Optionally, the second generation module 1004 is further configured to: perform feature encoding on the target style image using a target encoder to obtain a first encoding result, and perform feature encoding on the video frame images of the initial simulated video stream using the target encoder to obtain a second encoding result; perform feature fusion on the first encoding result and the second encoding result to obtain a fusion result; decode the fusion result using a target decoder to obtain a target image; and generate a target simulated video stream based on the target image. Optionally, the first acquisition module 1001 is further configured to: determine position information, posture information, and field of view angle information of the virtual camera; and acquire setting information based on the position information, posture information, and field of view angle information.Optionally, the screen information includes a target event type and an event duration of a target event described by a target simulation video stream to be generated, wherein the target event type includes an event type of perception failure in the real world, and the event duration is used to indicate the time length from the occurrence to the end of the target event. The first generation module 903 is further configured to: determine a video duration of the first simulation video based on the event duration, wherein the video duration is used to indicate the time length from the start to the end of the playback of the first simulation video, and the video duration is greater than or equal to the event duration; and render the video duration, setting information, and screen information using a digital twin system to obtain an initial simulation video stream. According to the disclosed embodiments, by adjusting the setting information of a virtual camera, the shooting angle of the virtual camera recording a virtual video stream in the digital twin world is adjusted, and the image information required to be produced in the target simulated video stream to be generated is determined. Then, based on the setting information and the image information, an initial simulated video stream is generated. Then, a target style image in the real world and the initial simulated video stream are input into a target video generation model. The initial simulated video stream is converted based on the target video generation model and the target style image to ultimately generate the target simulated video stream. The image style of the target simulated video stream corresponds to the image style of the target style image, thereby achieving the purpose of generating a target simulated video stream with specified image content and specified image style. This achieves the goal of generating a target simulated video stream with more realistic and reliable image content, more realistic image style, and applicable to any scenario. This also enriches the technical effects of the application scenarios in the digital twin world and solves the technical problem of low image fidelity in the virtual video stream. It should be noted that the first acquisition module 1001, determination module 1002, first generation module 1003, and second generation module 1004 described above correspond to steps S21 to S24 in Example 1. The examples and application scenarios implemented by these four modules and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above modules may be hardware components or software components stored in a memory and processed by one or more processors. These modules may also run on the server 10 provided in Example 1. According to an embodiment of the present disclosure, another embodiment of a device for implementing the above-described video stream generation method is also provided.FIG11 is a schematic structural diagram of another video stream generating apparatus according to Embodiment 5 of the present disclosure. As shown in FIG11 , the video stream generating apparatus 1100 includes: a first response module 1101, configured to respond to an adjustment instruction applied to an operation interface and display setting information of a virtual camera on the operation interface, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world simulating the real world, and the setting information is used to adjust the shooting angle of the virtual camera; a second response module 1102, configured to respond to an input instruction applied to the operation interface and display picture information of a target simulated video stream to be generated on the operation interface, wherein the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated; and a third response module 1103, configured to respond to a simulation instruction applied to the operation interface and display the target simulated video stream on the operation interface, wherein the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world, and the picture style of the target simulated video stream corresponds to the picture style of the target style image. The initial simulation video stream is generated based on the setting information and the picture information. According to an embodiment of the present disclosure, upon receiving an adjustment instruction on an operation interface, setting information of a virtual camera can be displayed on the operation interface, thereby adjusting the setting parameters of the virtual camera. The virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setting information is used to adjust the shooting angle of the virtual camera. Upon receiving an input instruction on the operation interface, image information of a target simulated video stream to be generated can be displayed on the operation interface. The image information records the simulated image elements to be produced in the target simulated video stream to be generated. Upon receiving a simulation instruction on the operation interface, the target simulated video stream can be displayed on the operation interface. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image. The target style image is a real image in the real world. The image style of the target simulated video stream corresponds to the image style of the target style image. The initial simulated video stream is generated based on the setting information and the image information. This achieves the purpose of generating a target simulated video stream with specified image content and specified image style, thereby achieving more realistic and reliable image content. The target simulation video stream has a more realistic picture style and can be applied to any scene. It also enriches the technical effects of the application scenarios in the digital twin world and solves the technical problem of low picture realism of virtual video streams.It should be noted that the first response module 1101, the second response module 1102, and the third response module 1103 correspond to steps S71 to S73 in Example 2. The examples and application scenarios implemented by these three modules and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above modules can be hardware components or software components stored in a memory and processed by one or more processors. The above modules can also run on the server 10 provided in Example 1. According to an embodiment of the present disclosure, another device embodiment for implementing the above-mentioned video stream generation method is also provided. FIG12 is a schematic structural diagram of another video stream generating apparatus according to Embodiment 5 of the present disclosure. As shown in FIG12 , the video stream generating apparatus 1200 includes: a second obtaining module 1201, configured to obtain a simulation video stream generation request through a first application programming interface; and a first returning module 1202, configured to return a simulation video stream generation response through a second application programming interface. The simulation video stream generation request includes: setting information of a virtual camera and picture information of a target simulation video stream to be generated; and the simulation video stream generation response includes: a target simulation video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; the picture information is used to record the simulation picture elements that need to be produced in the target simulation video stream to be generated; the target simulation video stream is generated by converting an initial simulation video stream based on a target video generation model and a target style image; the target style image is a real image in the real world; and the picture style of the target simulation video stream corresponds to the picture style of the target style image. The initial simulation video stream is generated based on the setting information and the picture information.In an embodiment of the present disclosure, a simulated video stream generation request is obtained through a first application programming interface (API). The simulated video stream generation request includes: virtual camera setup information and image information of a target simulated video stream to be generated. The virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world that simulates the real world. The setup information is used to adjust the shooting angle of the virtual camera. The image information is used to record the simulated image elements to be produced in the target simulated video stream to be generated. A simulated video stream generation response is then returned through a second application programming interface. The simulated video stream generation response includes: a target simulated video stream. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image. The target style image is a real image in the real world. The image style of the target simulated video stream corresponds to the image style of the target style image. The initial simulated video stream is generated based on the setup information and the image information. This achieves the purpose of generating a target simulated video stream with specified image content and specified image style, thereby achieving a target simulated video stream with more realistic and reliable image content, more realistic image style, and applicable to any scenario. This also enriches the technical effects of application scenarios in the digital twin world, thereby resolving the technical issue of low fidelity in virtual video streams. It should be noted that the second acquisition module 1201 and the first return module 1202 correspond to steps S81 and S82 in Example 3. The examples and application scenarios implemented by these two modules and the corresponding steps are the same, but are not limited to those disclosed in Example 1. It should be noted that the modules described above can be hardware or software components stored in memory and processed by one or more processors. These modules can also run on the server 10 provided in Example 1. According to an embodiment of the present disclosure, another device embodiment for implementing the aforementioned video stream generation method is also provided.FIG13 is a schematic structural diagram of another video stream generating apparatus according to Embodiment 5 of the present disclosure. As shown in FIG13 , the video stream generating apparatus 1300 includes: a third obtaining module 1301, configured to obtain an input simulated video stream generating dialogue request; a second returning module 1302, configured to return a simulated video stream generating dialogue reply in response to the simulated video stream generating dialogue request; wherein the simulated video stream generating dialogue request includes: setting information of a virtual camera and picture information in a target simulated video stream to be generated; and the simulated video stream generating dialogue reply includes: a target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world simulating the real world; the setting information is used to adjust the shooting angle of the virtual camera; and the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated. The target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world, and the picture style of the target simulated video stream corresponds to the picture style of the target style image. The initial simulation video stream is generated based on the setting information and the screen information; the display module 1303 is configured to display the target simulation video stream in the graphical user interface. According to an embodiment of the present disclosure, a simulated video stream generation dialog request is generated by obtaining an input simulated video stream, wherein the simulated video stream generation dialog request includes: virtual camera setting information and image information of a target simulated video stream to be generated. Then, in response to the simulated video stream generation dialog request, a simulated video stream generation dialog reply is returned, wherein the simulated video stream generation dialog reply includes: the target simulated video stream, wherein the virtual camera is used to record the virtual video stream in a digital twin world, which is a virtual world that simulates the real world; the setting information is used to adjust the shooting angle of the virtual camera; the image information is used to record the simulated image elements to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting the initial simulated video stream based on a target video generation model and a target style image, wherein the target style image is a real image in the real world; the image style of the target simulated video stream corresponds to the image style of the target style image, and the initial simulated video stream is generated based on the setting information and the image information. Finally, the target simulated video stream is displayed in a graphical user interface to provide feedback to the user, thereby achieving the purpose of generating a target simulated video stream with specified image content and specified image style. This enables the generation of target simulation video streams with more realistic and reliable content, more lifelike style, and applicability to any scenario. It also enriches the technical effects of application scenarios in the digital twin world and solves the technical problem of low realism in virtual video streams.It should be noted that the third acquisition module 1301, second return module 1302, and presentation module 1303 described above correspond to steps S91 to S93 in Example 4. The examples and application scenarios implemented by these three modules and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above modules can be hardware components or software components stored in a memory and processed by one or more processors. These modules can also run on the server 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above embodiments of the present disclosure are the same as the schemes, application scenarios, and implementation processes provided in Example 1, but are not limited to the schemes provided in Example 1. Example 6: The embodiments of the present disclosure may provide an electronic device, which may be any electronic device in a group of electronic devices. Optionally, in this embodiment, the electronic device may be replaced by a terminal device such as a mobile terminal. Optionally, in this embodiment, the electronic device may be located in at least one of multiple network devices in a computer network. In this embodiment, the electronic device may execute the steps of executing the protected program code in the video stream generation method provided in Example 1, but is not limited thereto. Alternatively, Figure 14 is a block diagram of an electronic device according to an embodiment of the present disclosure. As shown in Figure 14 , electronic device A may include one or more processors 1402 (only one is shown), a memory 1404, a storage controller, and a peripheral interface. The peripheral interface is connected to a radio frequency module, an audio module, and a display. The memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the video stream generation method and apparatus in the embodiments of the present disclosure. The processor executes the stored software programs and modules to perform various functional applications and data processing, thereby implementing the aforementioned video stream generation method. The memory may include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory located remotely from the processor, which may be connected to electronic device A via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.The processor can call information and applications stored in the memory through the transmission device to perform the following steps: adjusting setting information of a virtual camera, wherein the virtual camera is used to record a virtual video stream in a digital twin world, which is a virtual world simulating the real world, and the setting information is used to adjust the shooting angle of the virtual camera; determining image information in a target simulated video stream to be generated, wherein the image information is used to record simulated image elements that need to be produced in the target simulated video stream to be generated; generating an initial simulated video stream based on the setting information and the image information; and converting the initial simulated video stream based on a target video generation model and a target style image to generate a target simulated video stream, wherein the target style image is a real image in the real world, and the image style of the target simulated video stream corresponds to the image style of the target style image. Optionally, the processor can also perform the steps of executing each program code in each optional embodiment of Example 1, but is not limited thereto. As can be seen, the disclosed embodiments adjust the virtual camera's setting information to adjust the shooting angle of the virtual camera when recording a virtual video stream in the digital twin world, determine the image information required to record the target simulated video stream to be generated, and then generate an initial simulated video stream based on the setting information and the image information. The target style image in the real world and the initial simulated video stream are then input into a target video generation model. The initial simulated video stream is then converted based on the target video generation model and the target style image to ultimately generate the target simulated video stream. This ensures that the image style of the target simulated video stream corresponds to the image style of the target style image, thereby achieving the goal of generating a target simulated video stream with specified image content and a specified image style. This achieves the goal of generating a target simulated video stream with more realistic and reliable image content, a more realistic image style, and applicability to any scenario. This also enriches the technical effects of the application scenarios in the digital twin world and solves the technical problem of low image fidelity in the virtual video stream. Those skilled in the art will appreciate that the structure shown in FIG14 is merely illustrative. Electronic device A may also be a smartphone (e.g., an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG14 does not limit the structure of these electronic devices. For example, electronic device A may include more or fewer components (e.g., a network interface, a display device, etc.) than shown in FIG14 , or may have a configuration different from that shown in FIG14 .Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Embodiment 7: The embodiments of the present disclosure further provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the video stream generation method provided in the above embodiment 1. Optionally, in this embodiment, the computer-readable storage medium can be located in any electronic device in a group of electronic devices in a computer network, or in any mobile terminal in a group of mobile terminals. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the implementation steps included in Embodiment 1, but is not limited thereto. The embodiments of the present disclosure further provide a computer program product, comprising a computer program. When executed by a processor, the computer program implements any of the above-described video stream generation methods. The serial numbers of the above-mentioned embodiments of the present disclosure are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above-mentioned embodiments of the present disclosure, the descriptions of each embodiment are given with some emphasis. For portions not described in detail in one embodiment, reference should be made to the relevant descriptions of other embodiments. In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units described is merely a logical functional division. In actual implementation, other divisions may be employed. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed between each other may be through interfaces, or indirect coupling or communication connection between units or modules, and may be electrical or other forms. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the units may be selected to achieve the objectives of the present embodiments according to actual needs. In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.The aforementioned integrated units can be implemented in either hardware or software functional units. If implemented as software functional units and sold or used as standalone products, the integrated units can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to perform all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), removable hard drives, magnetic disks, or optical disks. The above description is merely a preferred embodiment of the present disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present disclosure, and such improvements and modifications should be considered within the scope of protection of the present disclosure.
Claims
Claims 1. A method for generating a video stream, comprising: Acquiring setting information of a virtual camera, wherein the virtual camera is used to record a virtual video stream in a digital twin world, the digital twin world being a virtual world simulating the real world, and the setting information is used to adjust the shooting angle of the virtual camera; determining picture information in a target simulated video stream to be generated, wherein the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated; generating an initial simulated video stream based on the setting information and the picture information; and converting the initial simulated video stream based on a target video generation model and a target style image to generate the target simulated video stream, wherein the target style image is a real image in the real world, and the picture style of the target simulated video stream corresponds to the picture style of the target style image.
2. The video stream generation method according to claim 1, wherein: Generating the initial simulated video stream based on the setting information and the picture information includes: generating a target tag based on the setting information and the picture information, wherein the target tag is used to identify video attributes of the target simulated video stream to be generated; rendering the setting information and the picture information using a digital twin system to obtain a first simulated video stream, wherein the digital twin system is used to construct the digital twin world; and marking the first simulated video stream using the target tag to obtain the initial simulated video stream.
3. The video stream generation method according to claim 1, wherein: The method also includes: obtaining a first image sample and a second image sample, wherein the first image sample is a real image sample in the real world, and the second image sample is a virtual image sample corresponding to the picture content of the first image sample; and training an initial video generation model based on the first image sample and the second image sample to obtain the target video generation model.
4. The video stream generation method according to claim 3, wherein: Acquiring the second image sample includes: performing image detection on the first image sample to obtain a first detection result, wherein the first detection result is used to represent a real picture element in the first image sample; Performing image simulation processing on the first detection result to generate the second image sample, wherein the simulated picture elements in the second image sample correspond to the real picture elements in the first image sample.
5. The video stream generation method according to claim 3, wherein: Acquiring the second image sample includes: performing image detection on the first image sample to obtain a second detection result, wherein the second detection result is used to represent the real event content described by the first image sample; and performing image simulation processing on the second detection result to generate the second image sample, wherein the simulated event content described by the second image sample corresponds to the real event content described by the first image sample.
6. The video stream generation method according to claim 3, wherein: The training of the initial video generation model based on the first image sample and the second image sample to obtain the target video generation model includes: performing image perspective conversion on the first image sample to generate a third image sample, wherein the third image sample is an image of the first image sample at a different perspective; performing image perspective conversion on the second image sample to generate a fourth image sample, wherein the fourth image sample is an image of the second image sample at a different perspective; and training the initial video generation model based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain the target video generation model.
7. The video stream generation method according to claim 6, wherein: The training of the initial video generation model based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain the target video generation model includes: training an initial encoder based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain a target encoder; training an initial decoder based on the first image sample and the third image sample to obtain a target decoder; and determining the target video generation model based on the target encoder and the target decoder.
8. The video stream generation method according to claim 7, wherein: The training of the initial encoder based on the first image sample, the second image sample, the third image sample, and the fourth image sample to obtain a target encoder includes: determining the first image sample and the third image sample as a first sample combination; and determining the second image sample and the fourth image sample as a second sample combination; The first image sample and the second image sample are determined as a third sample combination; and contrastive learning training is performed on the initial encoder based on a contrastive loss function, the first sample combination, the second sample combination, and the third sample combination to obtain the target encoder.
9. The video stream generation method according to claim 7, wherein: The training of the initial decoder based on the first image sample and the third image sample to obtain the target decoder includes: performing autoregressive training on the initial decoder based on the first image sample and the third image sample to obtain the target decoder.
10. The video stream generation method according to claim 7, wherein: The converting of the initial simulated video stream based on the target video generation model and the target style image to generate the target simulated video stream includes: performing feature encoding on the target style image using the target encoder to obtain a first encoding result, and performing feature encoding on the video frame image of the initial simulated video stream using the target encoder to obtain a second encoding result; performing feature fusion on the first encoding result and the second encoding result to obtain a fusion result; decoding the fusion result using the target decoder to obtain a target image; and generating the target simulated video stream based on the target image. The video stream generation method according to claim 1 , wherein: The setting information of the virtual camera includes: determining the position information, posture information and field of view angle information of the virtual camera; and acquiring the setting information based on the position information, the posture information and the field of view angle information.
12. The video stream generation method according to claim 2, wherein: The picture information includes a target event type of a target event described by the target simulation video stream to be generated and an event duration of the target event, wherein the target event type includes an event type of perception failure in the real world, and the event duration is used to indicate the time length from the occurrence to the end of the target event. The digital twin system is used to render the setting information and the picture information to obtain the first simulation video stream, including: determining the video duration of the first simulation video based on the event duration, wherein the video duration is used to indicate the time length from the start to the end of the playback of the first simulation video, and the video duration is greater than or equal to the event duration; and the video duration, the setting information, and the picture information are rendered using the digital twin system to obtain the initial simulation video stream.
13. A method for generating a video stream, comprising: Obtaining a simulation video stream generation request through a first application programming interface; A simulated video stream generation response is returned through a second application programming interface; wherein the simulated video stream generation request includes: setting information of a virtual camera and picture information in a target simulated video stream to be generated; and the simulated video stream generation response includes: a target simulated video stream, wherein the virtual camera is used to record a virtual video stream in a digital twin world, wherein the digital twin world is a virtual world that simulates the real world; the setting information is used to adjust a shooting angle of the virtual camera; the picture information is used to record simulated picture elements that need to be produced in the target simulated video stream to be generated; the target simulated video stream is generated by converting an initial simulated video stream based on a target video generation model and a target style image; the target style image is a real image in the real world; the picture style of the target simulated video stream corresponds to the picture style of the target style image; and the initial simulated video stream is generated based on the setting information and the picture information.
14. An electronic device, comprising: a memory storing an executable program; A processor is configured to run the program, wherein the program executes the video stream generating method according to any one of claims 1 to 13 when running.
15. A computer-readable storage medium comprising a stored executable program, wherein: When the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the video stream generation method according to any one of claims 1 to 13.
16. A computer program product, comprising a computer program, which, when executed by a processor, implements the video stream generation method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Vehicle neural network training
CN113177429A
Video generation method and device, computer equipment and storage medium
CN115761064A
Generating a virtual world to assess real-world video analysis performance
US20170243083A1