Scene reproduction method and system and electronic equipment
By generating case description text and reconstructing a three-dimensional scene model, the problem of evidence fragmentation in case scene reconstruction is solved, and visual reconstruction of the case scene is achieved, assisting judicial and criminal investigation analysis.
Patent Information
- Application Number
- CN202510835517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
Smart Images

Figure CN120707746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a scene reproduction method, system, electronic device and computer-readable storage medium. Background Art
[0002] Evidence collected at crime scenes is often fragmentary, including partial audio, video, and image footage of the crime scene, as well as textual materials compiled from confessions, eyewitness testimony, and case files. Because each form of evidence is fragmentary—incomplete snippets from a specific time period and location—it requires significant labor to compile. After investigators manually organize and compile the evidence, they form the case files and case descriptions provided to the court.
[0003] The case scene reconstruction technology is the real-life 3D panoramic investigation and case-handling assistance system, which integrates three-dimensional 360-degree panoramic photography, stereo modeling and mapping, human-computer interaction, and computer graphics simulation. It can not only use panoramic images to fully and realistically record and reproduce the scene, investigation direction, and on-site data measurement, but also can fully display the case scene from four dimensions: far air, near air, ground, and inside buildings.
[0004] This technology uses satellite remote sensing technology in the distance. According to the needs of case handling, satellite images are used to conduct macro-analysis and judgment of the case scene, which can directly reflect the distribution of landforms, rivers, buildings, etc. around the case scene. In the near distance, drones are used to take panoramic photos and stitch together the case scene. Through 3D panoramic modeling and VR technology, a 360-degree panoramic image of the scene is formed, which can be freely rotated and zoomed in and out at will, making it convenient for investigators to observe the whole picture of the case scene from different angles at high altitude. High-resolution panoramic cameras are used on the ground and inside buildings to efficiently and quickly take panoramic photos of the ground points of the case scene, ultimately forming a panoramic scene that includes multiple dimensions and can be jumped, zoomed in and out at will, presenting an intuitive, complete and dynamic case scene. Combined with VR glasses, people can have a sense of being at the scene.
[0005] However, because most of the first-hand scenes disappear or have been destroyed, the results of scene reconstruction are limited by incomplete detection information, and text, pictures, eyewitness testimony, etc. cannot be entered into the system, resulting in large differences between the reproduced scene and the crime scene, and some important information is hidden. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a method, system, electronic device and computer-readable storage medium for scene reproduction. The three-dimensional scene reconstruction model generated by the system is based on evidence materials such as text, image, audio, and video materials, and can realize the visual reproduction of the case scene. It can be used in judicial, criminal investigation and other fields to assist case analysis.
[0007] In a first aspect, an embodiment of the present invention provides a method for scene reproduction. The method comprises:
[0008] Obtaining evidence materials for the case and generating a case description text based on the evidence materials; the evidence materials may include one or more of text materials, image materials, audio materials, and video materials;
[0009] Obtain scenario analysis information based on case description text;
[0010] Generate images and videos based on scene analysis information;
[0011] Generate 3D scene reconstruction models based on images and videos.
[0012] In one embodiment, the method further comprises:
[0013] Convert the 3D scene reconstruction model into a 2D rendered image.
[0014] In one embodiment, generating a case description text based on the evidence material includes:
[0015] If the evidence material includes text data, the text data is preprocessed and then input into a pre-trained language processing model to obtain a first text summary;
[0016] If the evidence material includes image data, semantic analysis is performed on the visual content of the image data to obtain a first text, and then the first text is input into a pre-trained language processing model to obtain a second text summary;
[0017] If the evidence material includes audio material, converting the audio material into a second text, inputting the second text into a pre-trained language processing model to obtain a third text summary;
[0018] If the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material contains a subtitle file or a subtitled image, extract the subtitle file or the subtitled image and obtain a fourth text based on the subtitle file or the subtitled image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, the fourth text, and the fifth text according to the timeline to obtain a fused text; pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary;
[0019] The first text summary, the second text summary, the third text summary, and the fourth text summary are integrated to obtain a case description text.
[0020] In one embodiment, obtaining scenario analysis information based on the case description text includes:
[0021] Identify entities and events in the case description text, including time, place, and people; extract the relationship between entities and events, and determine the time and place of the event;
[0022] Generate environmental details based on case description text, identify characters' emotional tendencies, and generate characters' potential motivations;
[0023] Construct a structured dataset based on the case description text, entities, events, relationships between entities and events, the time and place of the event, environmental details, the emotional tendencies of the characters, and the potential motivations of the characters;
[0024] Generate scene information description text based on structured data sets;
[0025] Align entities and nodes in structured data, then jointly encode text and structured features and adapt them into text-structured features;
[0026] Generate scene analysis information based on case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets.
[0027] In one embodiment, generating images and videos based on scene analysis information includes:
[0028] Generate static images corresponding to each perspective based on scene analysis information;
[0029] Generate videos based on static images and scene analysis information.
[0030] In a second aspect, an embodiment of the present invention provides a system for scene reproduction, the system comprising:
[0031] A case description text generation module is used to obtain evidence materials of the case and generate a case description text based on the evidence materials; the evidence materials include one or more of text materials, image materials, audio materials, and video materials;
[0032] A scenario analysis information acquisition module is used to obtain scenario analysis information based on the case description text;
[0033] Image and video generation module, used to generate images and videos based on scene analysis information;
[0034] The 3D model generation module is used to generate a 3D scene reconstruction model based on images and videos.
[0035] In one embodiment, the system further comprises:
[0036] The two-dimensional image generation module is used to convert the three-dimensional scene reconstruction model into a two-dimensional rendering image.
[0037] In one embodiment, the case description text generation module includes:
[0038] a text processing submodule, configured to pre-process the text data if the evidence material includes text data, and then input the text data into a pre-trained language processing model to obtain a first text summary;
[0039] an image processing submodule for, if the evidence material includes image data, performing semantic analysis on the visual content of the image data to obtain a first text, and then inputting the first text into a pre-trained language processing model to obtain a second text summary;
[0040] an audio processing submodule, configured to convert the audio data into a second text if the evidence material includes audio data, and input the second text into a pre-trained language processing model to obtain a third text summary;
[0041] The video processing submodule is configured to, if the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material includes a subtitle file or a subtitled image, extract the subtitle file or the subtitled image and obtain a fourth text based on the subtitle file or the subtitled image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, the fourth text, and the fifth text along a timeline to obtain a fused text; and pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary;
[0042] The integration submodule is used to integrate the obtained first text summary, second text summary, third text summary, and fourth text summary to obtain a case description text.
[0043] In one embodiment, the scene analysis information acquisition module includes:
[0044] The recognition submodule is used to identify entities and events in the case description text, including time, place, and people; extract the relationship between entities and events, and determine the time and place of the event;
[0045] The environmental detail generation submodule is used to generate environmental detail information based on the case description text, identify the emotional tendencies of the characters, and generate the characters' potential motivations;
[0046] A structured submodule is used to construct a structured dataset based on the case description text, entities, events, relationships between entities and events, the time and place of the event, environmental details, the emotional tendencies of the characters, and the potential motivations of the characters;
[0047] A scene information description text generation submodule is used to generate scene information description text based on a structured data set;
[0048] The feature joint submodule is used to align entities and nodes in structured data, and then jointly encode text and structured features and adapt them into text-structured features;
[0049] The scene analysis submodule is used to generate scene analysis information based on case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets.
[0050] In one embodiment, the image and video generation module includes:
[0051] An image generation submodule, configured to generate static images corresponding to each viewing angle based on scene analysis information;
[0052] The video generation submodule is used to generate videos based on static images and scene analysis information.
[0053] In a third aspect, the present invention provides an electronic device, comprising:
[0054] at least one processor; and
[0055] a memory communicatively connected to the at least one processor; wherein,
[0056] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the above-mentioned scene reproduction method.
[0057] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement the above-mentioned scene reproduction method when executed.
[0058] In an embodiment of the present invention, a case description is generated based on evidence materials collected from the case; the evidence materials include one or more of text, images, audio, and video. Scene analysis information is obtained based on the case description text; images and videos are generated based on the scene analysis information; and a three-dimensional scene reconstruction model is generated based on the images and videos. The generated three-dimensional scene reconstruction model, based on the text, images, audio, and video evidence materials, can visualize the case scene and can be used to assist in case analysis in fields such as judicial and criminal investigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 This is a flow chart of a scene reproduction method provided by the first embodiment of the present invention;
[0061] Figure 2 This is a structural diagram of a scene reproduction system provided by Embodiment 2 of the present invention;
[0062] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0063] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0064] Figure 1 This is a flow chart of a scene reproduction method provided in the first embodiment of the present invention. This embodiment can be applied to a scene reproduction system, such as Figure 1 As shown, this embodiment may include the following steps:
[0065] Step 101: Obtain evidence materials of the case and generate a case description text based on the evidence materials. The evidence materials include one or more of text materials, image materials, audio materials, and video materials.
[0066] In this step, the scene reconstruction system can receive case evidence from users (case handlers) via an API. This evidence can include text recordings of the case parties and witnesses' verbal accounts of the incident, images of physical evidence and the crime scene, audio recordings of the original crime scene recordings, and video footage of the crime scene and surroundings captured by investigators.
[0067] After obtaining the evidence materials of the case, a case description text can be generated based on the evidence materials, that is, the evidence materials can be converted into a concise and clear case description text, so as to achieve the purpose of preliminary extraction and organization of information.
[0068] Generating a case description text based on the evidence material may include the following steps:
[0069] In sub-step 101a, if the evidence material includes text data, the text data is pre-processed and then input into a pre-trained language processing model to obtain a first text summary.
[0070] In this step, after removing redundant symbols (such as line breaks and HTML tags) from the document, the document is input into a pre-trained language processing model to obtain a first text summary. The pre-trained language processing model can be BERT or DEEPSEEK-V3, or a combination of BERT and DEEPSEEK-V3.
[0071] Sub-step 101b: If the evidence material includes image data, semantic analysis is performed on the visual content of the image data to obtain a first text, and then the first text is input into a pre-trained language processing model to obtain a second text summary.
[0072] In this step, a semantic analysis of the visual content of the image data can be performed using deep learning methods to obtain a first text that describes information such as the image's content, scene, objects, and their relationships. After obtaining the first text, the first text is input into a pretrained language processing model to obtain a second text summary. Similarly, the pretrained language processing model can be BERT or DEEPSEEK-V3, or a combination of BERT and DEEPSEEK-V3. Of course, other suitable models can also be used, which are not listed here.
[0073] In sub-step 101c, if the evidence material includes audio data, the audio data is converted into a second text, and the second text is input into a pre-trained language processing model to obtain a third text summary.
[0074] In this step, the audio material may be converted into the second text in the following ways:
[0075] Method 1: Input the audio data directly into the Whisper model to obtain the second text.
[0076] Method 2: Convert the audio to mono, standardize the audio format (such as adjusting the sampling rate to 16kHz), and perform other preprocessing. Then, use MFCC (Mel-Frequency Cepstral Coefficient) to extract audio features, and input the audio features into the DeepSpeech model to obtain the second text.
[0077] After converting the audio material into the second text, the language model can be used to correct misrecognized words. The corrected second text is input into the pre-trained language processing model to obtain a third text summary.
[0078] Sub-step 101d: If the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material contains a subtitle file or a subtitle image, extract the subtitle file or the subtitle image and obtain a fourth text based on the subtitle file or the subtitle image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, fourth text, and fifth text along the timeline to obtain a fused text; pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary.
[0079] In this step, a video processing tool such as FFmpeg may be used to extract the audio content of the video material, and then the audio content is converted into a third text.
[0080] If the video material contains a subtitle file (such as SRT or VTT format), the subtitle file can be directly converted into the fourth text. If the video material does not contain a subtitle file, the key frames in the video are extracted and the text in the key frames is recognized using optical character recognition (OCR) or deep learning methods to obtain the fourth text.
[0081] Multimodal models (such as CLIP and BLIP) can be used to perform semantic analysis on the visual content of video materials to obtain a fifth text that records information such as the content, scenes, objects and their relationships in the video.
[0082] Then, the third text, the fourth text, and the fifth text are integrated according to the timeline to obtain a fused text with coherent content.
[0083] After obtaining the fused text, the fused text is preprocessed and then input into a pre-trained language processing model to obtain a fourth text summary.
[0084] Sub-step 101e: Integrate the first text summary, the second text summary, the third text summary, and the fourth text summary to obtain the case description text.
[0085] Step 102: Acquire scenario analysis information based on the case description text.
[0086] In this step, after obtaining the case description text, scene analysis information can be obtained based on the case description text, that is, case details can be extracted and deduced through the case description text, and these details can be integrated into structured data to output scene analysis information, providing accurate information support for the subsequent generation of images and videos, as well as the construction of case scene point cloud models.
[0087] Step 102 may include the following steps:
[0088] Sub-step 102a, identifying entities and events in the case description text, wherein the entities include time, place and people; extracting the relationship between the entities and events, and determining the time and place of the events.
[0089] Named Entity Recognition (NER) can be used to identify entities with specific meanings from case descriptions, such as time, place, and person entities that are highly relevant to the case, and classify them into predefined categories. For example, a large language model (LLM) can be used to identify entities in case descriptions.
[0090] Events can be understood as objective facts driven by action / state verbs, encompassing factors such as participants (entities), time, and location. Event identification in case descriptions can be achieved using TF-IDF or TextRank combined with verb phrases. This combination of TF-IDF or TextRank allows for lightweight event identification. Of course, other methods are also available for event identification in case descriptions, which are not listed here.
[0091] After identifying events in the case description text, the meaning of words and sentences can also be understood through context (such as using BERT, DEEPSEEK-V3, ELMo) to resolve polysemy and semantic ambiguity problems.
[0092] After identifying entities and events in the case description text, the associations between events and entities are extracted, the time and location of the event are determined, and the event structure is improved. For example, the combination of "BERT model + relationship classification head" can be used to implement entity and event relationship analysis. BERT, as a bidirectional Transformer pre-trained model, can capture the dynamic semantics of words in different contexts. Based on BERT pre-training, a customized relationship classification layer (such as a fully connected layer + Softmax) is added to map the semantic vector output by BERT to the relationship between entities and events.
[0093] Sub-step 102b, generating environmental detail information based on the case description text, identifying the character's emotional tendencies, and generating the character's potential motivations.
[0094] In this step, after the case description text is segmented, part-of-speech tagged, and the core verb phrases are extracted as event trigger words, the trigger word semantics can be determined in combination with the context or named entity recognition (NER), and the trigger word is converted into a word vector (such as Word2Vec); semantically similar event-relationship pairs are retrieved in ATOMIC. If there is a direct match, the environmental attributes are directly extracted from ATOMIC; otherwise, semantic similarity expansion is performed; then the event-relationship pairs are mapped and converted according to pre-set environmental rules to associate them with environmental details, and finally the environmental detail information is obtained.
[0095] A person's emotional tendencies can be identified using deep learning methods. For example, the preprocessed case description text is input into a pre-trained emotion recognition model (CNN, RNN, or RNN variant can be selected) to output the person's emotional tendencies.
[0096] The potential motivations of characters can be generated using DEEPSEEK-V3. Specifically, the case description text and the task instructions for generating hypothetical motivations are input into a platform or interface that interacts with DEEPSEEK-V3 to obtain the potential motivations of characters generated by DEEPSEEK-V3.
[0097] Sub-step 102c: constructing a structured data set based on the case description text, entities, events, relationships between entities and events, time and place of events, environmental details, emotional tendencies of characters, and potential motivations of characters.
[0098] Specifically, the case description text, entities, events, relationships between entities and events, time and place of events, environmental details, emotional tendencies of characters, and potential motivations of characters can be constructed into structured graph data consisting of nodes, relationships and attributes according to the model specifications of the Neo4j graph database.
[0099] Sub-step 102d: Generate scene information description text based on the structured data set.
[0100] For the case where the structured dataset constructed in sub-step 102c is constructed into structured graph data consisting of nodes, relationships and attributes according to the model specifications of the Neo4j graph database, the generation of scene information description text based on the structured dataset can specifically be to convert the graph data into prompts (input text or instructions used to guide the model to generate specific outputs), and input them into DEEPSEEK-V3 to generate more fluent scene information description text.
[0101] Sub-step 102e, aligning entities and nodes in structured data, and then jointly encoding text and structured features and adapting them into text-structured features.
[0102] In this step, the knowledge graph embedding model (such as TransE and RotatE) can be used to align entities and nodes in structured data, and then the ViLBERT model is used to jointly encode text and structured features. Finally, CLIP is used to adapt the joint features to text-structured features.
[0103] Sub-step 102f, generating scene analysis information based on the case description text, scene information description text, character's emotional tendencies, environmental detail information, and structured data set.
[0104] In this step, the case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets can be used by DEEPSEEK-V3 to generate scene-related semantic descriptions. DALL·E then converts the semantics into visual scenes, and finally combines the two to output structured scene analysis information.
[0105] Step 103: Generate images and videos based on the scene analysis information.
[0106] In this step, the scene analysis information can be input into an image generation model (such as DALL E3, StableDiffusion, MidJourney, etc.) to generate static images corresponding to each perspective. Then, a video generation model (such as VideoComposer) is used to generate a video based on the static images and scene analysis information. Frame interpolation techniques (such as DAIN and FILM) are used to generate transition frames between images to make the video smoother.
[0107] Step 104: Generate a 3D scene reconstruction model based on the image and video.
[0108] In this step, the image and video obtained in step 103 can be input into 3D Gaussian Splatting (3DGS) to generate a 3D scene reconstruction model. In the process of 3DGS generating a 3D scene reconstruction model, the multi-view can also be enhanced by a video diffusion model (such as Video LDMs). Figure 1 Consistency ensures the continuity of dynamic point clouds.
[0109] By importing the obtained three-dimensional scene reconstruction model into VR or AR equipment, the case scene can be reproduced.
[0110] In one embodiment, after obtaining the 3D scene reconstruction model, the 3D scene reconstruction model may be converted into a 2D rendering image. For example, viewing software (player) may be used to convert the 3D scene reconstruction model into a 2D rendering image.
[0111] In this embodiment, a case description is generated based on evidence collected from the case; the evidence includes one or more of text, images, audio, and video. Scene analysis information is obtained based on the case description; images and videos are generated based on the scene analysis information; and a three-dimensional scene reconstruction model is generated based on the images and videos. The generated three-dimensional scene reconstruction model, based on the text, images, audio, and video evidence, can visualize the case scene and can be used to assist in case analysis in judicial and criminal investigation fields.
[0112] Corresponding to the scene reproduction method of the present invention, the present invention also provides a scene reproduction system, Figure 2 This is a schematic diagram of the system structure for reproducing a scenario. Figure 2 As shown, the system reproducing this scenario includes:
[0113] The case description text generation module 201 is used to obtain evidence materials of the case and generate a case description text based on the evidence materials; the evidence materials include one or more of text materials, image materials, audio materials, and video materials;
[0114] A scenario analysis information acquisition module 202 is used to acquire scenario analysis information based on the case description text;
[0115] An image and video generation module 203 is configured to generate images and videos based on scene analysis information;
[0116] The 3D model generation module 204 is configured to generate a 3D scene reconstruction model based on images and videos.
[0117] In one embodiment, the system further comprises:
[0118] The two-dimensional image generation module is used to convert the three-dimensional scene reconstruction model into a two-dimensional rendering image.
[0119] In one embodiment, the case description text generation module 201 includes:
[0120] a text processing submodule, configured to pre-process the text data if the evidence material includes text data, and then input the text data into a pre-trained language processing model to obtain a first text summary;
[0121] an image processing submodule for, if the evidence material includes image data, performing semantic analysis on the visual content of the image data to obtain a first text, and then inputting the first text into a pre-trained language processing model to obtain a second text summary;
[0122] an audio processing submodule, configured to convert the audio data into a second text if the evidence material includes audio data, and input the second text into a pre-trained language processing model to obtain a third text summary;
[0123] The video processing submodule is configured to, if the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material includes a subtitle file or a subtitled image, extract the subtitle file or the subtitled image and obtain a fourth text based on the subtitle file or the subtitled image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, the fourth text, and the fifth text along a timeline to obtain a fused text; and pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary;
[0124] The integration submodule is used to integrate the obtained first text summary, second text summary, third text summary, and fourth text summary to obtain a case description text.
[0125] In one embodiment, the scene analysis information acquisition module 202 includes:
[0126] The recognition submodule is used to identify entities and events in the case description text, including time, place, and people; extract the relationship between entities and events, and determine the time and place of the event;
[0127] The environmental detail generation submodule is used to generate environmental detail information based on the case description text, identify the emotional tendencies of the characters, and generate the characters' potential motivations;
[0128] A structured submodule is used to construct a structured dataset based on the case description text, entities, events, relationships between entities and events, the time and place of the event, environmental details, the emotional tendencies of the characters, and the potential motivations of the characters;
[0129] A scene information description text generation submodule is used to generate scene information description text based on a structured data set;
[0130] The feature joint submodule is used to align entities and nodes in structured data, and then jointly encode text and structured features and adapt them into text-structured features;
[0131] The scene analysis submodule is used to generate scene analysis information based on case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets.
[0132] In one embodiment, the image and video generation module 203 includes:
[0133] An image generation submodule, configured to generate static images corresponding to each viewing angle based on scene analysis information;
[0134] The video generation submodule is used to generate videos based on static images and scene analysis information.
[0135] A scene reproduction system provided by an embodiment of the present invention can execute a scene reproduction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.
[0136] Figure 3 A schematic block diagram of an electronic device 30 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0137] like Figure 3 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc., which is communicatively connected to the at least one processor 31. The memory stores a computer program that can be executed by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. Various programs and data required for the operation of the electronic device 30 can also be stored in the RAM 33. The processor 31, ROM 32, and RAM 33 are connected to each other via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.
[0138] Multiple components in the electronic device 30 are connected to the I / O interface 35, including an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a magnetic disk, an optical disk, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0139] The processor 31 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 31 executes the various methods and processes described above, such as the scene reconstruction method.
[0140] In some embodiments, the method for scene reproduction can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded into the RAM 33 and executed by the processor 31, one or more steps of the method for scene reproduction described above can be performed. Alternatively, in other embodiments, the processor 31 can be configured to perform the method for scene reproduction by any other appropriate means (for example, by means of firmware).
[0141] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0142] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0143] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0145] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0146] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0147] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0148] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for scene reproduction, characterized in that: The method comprises: Obtaining evidence materials for the case and generating a case description text based on the evidence materials; the evidence materials may include one or more of text materials, image materials, audio materials, and video materials; Obtain scenario analysis information based on case description text; Generate images and videos based on scene analysis information; Generate 3D scene reconstruction models based on images and videos.
2. The method according to claim 1, characterized in that The generating of the case description text based on the evidence materials includes: If the evidence material includes text data, the text data is preprocessed and then input into a pre-trained language processing model to obtain a first text summary; If the evidence material includes image data, semantic analysis is performed on the visual content of the image data to obtain a first text, and then the first text is input into a pre-trained language processing model to obtain a second text summary; If the evidence material includes audio material, converting the audio material into a second text, inputting the second text into a pre-trained language processing model to obtain a third text summary; If the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material contains a subtitle file or a subtitled image, extract the subtitle file or the subtitled image and obtain a fourth text based on the subtitle file or the subtitled image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, the fourth text, and the fifth text according to the timeline to obtain a fused text; pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary; The first text summary, the second text summary, the third text summary, and the fourth text summary are integrated to obtain a case description text.
3. The method according to any one of claims 1 or 2, characterized in that The obtaining of scenario analysis information based on the case description text includes: Identify entities and events in the case description text, including time, place, and people; extract the relationship between entities and events, and determine the time and place of the event; Generate environmental details based on case description text, identify characters' emotional tendencies, and generate characters' potential motivations; Construct a structured dataset based on the case description text, entities, events, relationships between entities and events, the time and place of the event, environmental details, the emotional tendencies of the characters, and the potential motivations of the characters; Generate scene information description text based on structured data sets; Align entities and nodes in structured data, then jointly encode text and structured features and adapt them into text-structured features; Generate scene analysis information based on case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets.
4. The method according to claim 3, characterized in that The generating of images and videos based on scene analysis information includes: Generate static images corresponding to each perspective based on scene analysis information; Generate videos based on static images and scene analysis information.
5. A scene reproduction system, characterized in that: The system includes: A case description text generation module is used to obtain evidence materials of the case and generate a case description text based on the evidence materials; the evidence materials include one or more of text materials, image materials, audio materials, and video materials; A scenario analysis information acquisition module is used to obtain scenario analysis information based on the case description text; Image and video generation module, used to generate images and videos based on scene analysis information; The 3D model generation module is used to generate a 3D scene reconstruction model based on images and videos.
6. The system according to claim 5, characterized in that The case description text generation module includes: a text processing submodule, configured to pre-process the text data if the evidence material includes text data, and then input the text data into a pre-trained language processing model to obtain a first text summary; an image processing submodule for, if the evidence material includes image data, performing semantic analysis on the visual content of the image data to obtain a first text, and then inputting the first text into a pre-trained language processing model to obtain a second text summary; an audio processing submodule, configured to convert the audio data into a second text if the evidence material includes audio data, and input the second text into a pre-trained language processing model to obtain a third text summary; The video processing submodule is configured to, if the evidence material includes video content, extract the audio content of the video material and convert the audio content into a third text; if the video material includes a subtitle file or a subtitled image, extract the subtitle file or the subtitled image and obtain a fourth text based on the subtitle file or the subtitled image; perform semantic analysis on the visual content of the video material to obtain a fifth text; integrate the third text, the fourth text, and the fifth text along a timeline to obtain a fused text; and pre-process the fused text and input it into a pre-trained language processing model to obtain a fourth text summary; The integration submodule is used to integrate the obtained first text summary, second text summary, third text summary, and fourth text summary to obtain a case description text.
7. The system according to claim 5 or 6, characterized in that The scene analysis information acquisition module includes: The recognition submodule is used to identify entities and events in the case description text, including time, place, and people; extract the relationship between entities and events, and determine the time and place of the event; The environmental detail generation submodule is used to generate environmental detail information based on the case description text, identify the emotional tendencies of the characters, and generate the characters' potential motivations; A structured submodule is used to construct a structured dataset based on the case description text, entities, events, relationships between entities and events, the time and place of the event, environmental details, the emotional tendencies of the characters, and the potential motivations of the characters; A scene information description text generation submodule is used to generate scene information description text based on a structured data set; The feature joint submodule is used to align entities and nodes in structured data, and then jointly encode text and structured features and adapt them into text-structured features; The scene analysis submodule is used to generate scene analysis information based on case description text, scene information description text, character emotional tendencies, environmental details, and structured data sets.
8. The system according to claim 7, characterized in that The image and video generation module includes: An image generation submodule, configured to generate static images corresponding to each viewing angle based on scene analysis information; The video generation submodule is used to generate videos based on static images and scene analysis information.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the scene reproduction method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the scene reproduction method according to any one of claims 1 to 5 when executed.