Analysis report generation method and device, electronic equipment and readable storage medium

By combining portable AI analysis devices with multimodal large models and large language models, multidimensional analysis and automated generation of criminal investigation case reports have been achieved, solving the problems of low report comprehensiveness and efficiency in existing technologies and improving the quality and speed of report generation.

CN121963201APending Publication Date: 2026-05-01GUANGZHOU CRIMINAL SCIENCE & TECHNOLOGY RESEARCH INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU CRIMINAL SCIENCE & TECHNOLOGY RESEARCH INSTITUTE
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies generate reports with low comprehensiveness and limited processing dimensions in criminal case analysis, and the efficiency of report generation is low, requiring manual intervention.

Method used

Portable AI analysis devices are used to input the images and/or videos involved in the case into a multimodal large model for multi-dimensional analysis, generating object labels and descriptive information, which are then input into a large language model to directly generate a target report.

Benefits of technology

It improves the comprehensiveness and efficiency of report generation, reduces manual intervention, and increases the speed and quality of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963201A_ABST
    Figure CN121963201A_ABST
Patent Text Reader

Abstract

The invention provides an analysis report generation method and device, electronic equipment and a readable storage medium, and the method comprises the steps: inputting a processed case-related image and / or a case-related video and metadata information into a pre-trained multi-modal large model, so as to enable the multi-modal large model to recognize a target object contained in the case-related image and / or the case-related video; extracting context and scene information based on text information in the case-related image and / or the case-related video, and outputting at least one object tag and description information for the case-related image and / or the case-related video; inputting the at least one object label and the description information into a pre-trained big language model, so that the big language model performs structured processing on the object label and the description information, and correspondingly filling obtained structured data into a report template matched with the selected case type, and outputting a target report corresponding to the case-related image and / or the case-related video. In this way, the comprehensiveness and efficiency of report generation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, electronic devices, and readable storage media for generating analysis reports Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to methods, apparatus, electronic devices, and readable storage media for generating analysis reports. Background Technology

[0002] In the process of criminal investigation case analysis, in order to ensure the traceability of subsequent cases and the integrity of evidence preservation, a corresponding case report is generally generated based on the information obtained at the crime scene. In the existing technology, tools can be used to transcribe and organize the dialogue content in the recorded audio into a structured report draft using variants of large language models such as OpenAI's ChatGPT. Staff members then review, modify, and supplement the report draft to generate a formal report. However, this method can only process the acquired audio data, has a single processing dimension, and the generated report has low comprehensiveness. Furthermore, after the report is generated, staff members still need to process it, which also results in low report generation efficiency. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, apparatus, electronic device and readable storage medium for generating analysis reports. By using a portable AI analysis device to input the extracted images and / or videos involved in the case into a multimodal large model, the images and / or videos involved in the case are analyzed from multiple dimensions to obtain at least one object label and descriptive information. Then, the at least one object label and descriptive information are input into a large language model, and the corresponding target report is directly generated through processing in the large language model, which can improve the comprehensiveness and efficiency of report generation.

[0004] In a first aspect, embodiments of this application provide a method for generating an analysis report, applied to a portable AI analysis device. The method includes: receiving extracted images and / or videos related to a case; preprocessing the images and / or videos to determine metadata information corresponding to the images and / or videos; inputting the processed images and / or videos and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify target objects contained in the images and / or videos, and extract context and scene information based on text information in the images and / or videos, outputting at least one object label and description information for the images and / or videos; inputting at least one object label and description information into a pre-trained large language model, so that the large language model can perform structured processing on the object labels and description information, and filling the obtained structured data into a report template matching the selected case type, outputting a target report corresponding to the images and / or videos.

[0005] In one possible implementation, the step of inputting the processed images and / or videos involved in the case, along with metadata information, into a pre-trained multimodal large model, so that the multimodal large model can identify target objects contained in the images and / or videos involved in the case, and extract context and scene information based on text information in the images and / or videos involved in the case, and output at least one object label and descriptive information for the images and / or videos involved in the case, includes: inputting the images and / or videos involved in the case into the pre-trained multimodal large model, so that the multimodal large model can visually encode the images and / or videos involved in the case to obtain high-dimensional feature vectors, and decoding the high-dimensional feature vectors to determine... The target object contained in the images and / or videos involved in the case; identifying the text information contained in the images and / or videos involved in the case, analyzing the text information according to criminal investigation terminology, determining at least one professional term and target sensitive words contained in the text information, and generating object tags based on at least one professional term and target sensitive words; when the target object is a preset identification object, analyzing the object behavior of the target object and the context information of the scene where the target object is located according to preset dangerous behavior tags, generating object tags and scene information; based on the identified target object, object tags and scene information, outputting at least one object tag and descriptive information for the images and / or videos involved in the case.

[0006] In one possible implementation, the descriptive information is generated through the following steps: based on the identified target object, object tag, and scene information, clue analysis information is generated according to preset logical reasoning rules and a knowledge graph in the criminal investigation field; based on the identified target object, object tag, and scene information, special event analysis is performed to generate special event analysis information; and based on the clue analysis information and the special event analysis information, the descriptive information is generated.

[0007] In one possible implementation, before the large language model performs structured processing on the object labels and description information, the generation method further includes: inputting at least one object label and description information into a pre-trained large language model, and filtering and deduplicating at least one object label and description information according to preset label weights and importance information of each object label to obtain at least one processed object label and description information, so as to perform structured processing on at least one processed object label and description information.

[0008] In one possible implementation, the step of inputting at least one of the object labels and description information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and description information, includes: inputting at least one of the object labels and description information into a pre-trained large language model, so that the large language model determines the hierarchy and association between each of the object labels and description information according to the case logic relationship, and generates structured information based on at least one of the object labels and description information and the hierarchy and association between each of the object labels and description information.

[0009] In one possible implementation, the step of filling the obtained structured data into a report template matching the selected case type and outputting a target report corresponding to the case-related image and / or video includes: aggregating and classifying the structured data according to the analysis framework and evidence classification rules corresponding to the case type to determine the classified type structured data; filling the determined type structured data and the corresponding case-related image and / or video into the target area of ​​the report template according to a preset display format to obtain an initial report; generating a case information summary based on the data filled in each target area of ​​the initial report; and outputting a target report corresponding to the case-related image and / or video based on the initial report and the case information summary.

[0010] In one possible implementation, the preprocessing of the image and / or video in question to determine the metadata information corresponding to the image and / or video in question includes: for the image in question, performing size normalization, format conversion, and quality assessment on the image in question, and extracting the metadata of the image in question; wherein, the quality assessment is to evaluate the security and integrity of the image in question; for the video in question, performing frame extraction processing on the video in question according to a preset frame rate, dividing the video in question into at least one image in question, performing size normalization, format conversion, and quality assessment on each image in question, and extracting the metadata of the image in question.

[0011] Secondly, this application also provides an analysis report generation device for portable AI analysis devices. The generation device includes: a data preprocessing module for receiving extracted images and / or videos related to a case, preprocessing the images and / or videos related to the case, and determining the metadata information corresponding to the images and / or videos related to the case; and a tag and description information extraction module for inputting the processed images and / or videos related to the case and the metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the contents contained in the images and / or videos related to the case. The system includes a target object and extracts context and scene information based on text information in the images and / or videos involved in the case, outputting at least one object label and descriptive information for the images and / or videos involved in the case; a target report generation module is used to input at least one of the object labels and descriptive information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and descriptive information, and fills the obtained structured data into a report template matching the selected case type, outputting a target report corresponding to the images and / or videos involved in the case.

[0012] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the analysis report generation method as described in any of the first aspects.

[0013] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the analysis report generation method as described in any of the first aspects.

[0014] The analysis report generation method, apparatus, electronic device, and readable storage medium provided in this application embodiment receive extracted case-related images and / or case-related videos; preprocess the case-related images and / or case-related videos to determine the metadata information corresponding to the case-related images and / or case-related videos; input the processed case-related images and / or case-related videos and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify target objects contained in the case-related images and / or case-related videos, and extract context and scene information based on the text information in the case-related images and / or case-related videos, and output at least one object label and descriptive information for the case-related images and / or case-related videos; input at least one object label and descriptive information into a pre-trained large language model, so that the large language model can perform structured processing on the object label and descriptive information, and fill the obtained structured data into a report template matching the selected case type, and output a target report corresponding to the case-related images and / or case-related videos. In this way, by using a portable AI analysis device to input the extracted images and / or videos related to the case into a multimodal big model, the images and / or videos related to the case are analyzed from multiple dimensions to obtain at least one object label and descriptive information. Then, the at least one object label and descriptive information are input into a big language model, and the corresponding target report is directly generated through the processing in the big language model, which can improve the comprehensiveness and efficiency of report generation.

[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 is a flowchart of an analysis report generation method provided in an embodiment of this application; Figure 2 is a schematic diagram of the structure of an analysis report generation device provided in an embodiment of this application; Figure 3 is a schematic diagram of the structure of an analysis report generation device provided in an embodiment of this application; Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0019] First, the applicable scenarios for this application will be introduced. This application can be applied to the field of data analysis technology.

[0020] In the process of analyzing criminal cases, in order to ensure the traceability of subsequent cases and the integrity of evidence preservation, a corresponding case report is generally generated based on the information obtained at the crime scene.

[0021] In existing technologies, tools can be used to transcribe and organize the dialogue content in recorded audio into a structured report draft using variants of large language models such as OpenAI's ChatGPT. Staff can then review, modify, and supplement the report draft to generate a formal report. However, this method can only process the acquired audio data, has a single processing dimension, and generates a report with low comprehensiveness. Furthermore, after the report is generated, staff still need to process it, which also results in low report generation efficiency.

[0022] Based on this, embodiments of this application provide a method for generating analysis reports to improve the comprehensiveness and efficiency of report generation.

[0023] Please refer to Figure 1, which is a flowchart of an analysis report generation method provided in an embodiment of this application. As shown in Figure 1, the analysis report generation method provided in this embodiment includes: S101, receiving extracted images and / or videos related to the case, preprocessing the images and / or videos related to the case, and determining the metadata information corresponding to the images and / or videos related to the case.

[0024] S102. Input the processed images and / or videos involved in the case and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the target objects contained in the images and / or videos involved in the case, and extract context and scene information based on the text information in the images and / or videos involved in the case, and output at least one object label and description information for the images and / or videos involved in the case.

[0025] S103. Input at least one of the object labels and description information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and description information, and fills the obtained structured data into a report template that matches the selected case type, and outputs a target report corresponding to the case-related image and / or the case-related video.

[0026] The method for generating analysis reports provided in this application involves inputting extracted images and / or videos related to the case into a multimodal large model using a portable AI analysis device. This allows for multi-dimensional analysis of the images and / or videos to obtain at least one object label and descriptive information. The at least one object label and descriptive information are then input into a large language model, and the corresponding target report is directly generated through processing within the large language model. This approach can improve the comprehensiveness and efficiency of report generation.

[0027] The following describes the exemplary steps of the embodiments of this application: S101, receiving the extracted images and / or videos involved in the case, preprocessing the images and / or videos involved in the case, and determining the metadata information corresponding to the images and / or videos involved in the case.

[0028] In the process of criminal investigation case analysis, in order to ensure the traceability of subsequent cases and the integrity of evidence preservation, a corresponding case report is generally generated based on the information obtained at the crime scene. In the existing technology, tools can be used to transcribe and organize the dialogue content in the recorded audio into a structured report draft using variants of large language models such as OpenAI's ChatGPT. Staff members then review, modify, and supplement the report draft to generate a formal report. However, this method can only process the acquired audio data, has a single processing dimension, and the generated report has low comprehensiveness. Furthermore, after the report is generated, staff members still need to process it, which also results in low report generation efficiency.

[0029] Based on this, in this embodiment of the application, the extracted images and / or videos involved in the case are input into a multimodal large model through a portable AI analysis device. The images and / or videos involved in the case are analyzed in multiple dimensions to obtain at least one object label and descriptive information. Then, the at least one object label and descriptive information are input into a large language model. The corresponding target report is directly generated through the processing in the large language model, which can improve the comprehensiveness and efficiency of report generation.

[0030] In one possible implementation, the portable AI analysis device integrates a high-performance computing module, interface, power management and heat dissipation system as independent hardware devices. In order to adapt to the special conditions of crime scenes in the field of criminal investigation, it is made of robust and durable materials (such as aluminum alloy or high-strength engineering plastics) and has the characteristics of shockproof, dustproof and splashproof, so as to adapt to various complex and even harsh working environments.

[0031] Specifically, the size and weight of the portable AI analytics device are kept within a reasonable range to ensure easy carrying by a single person. In terms of interfaces, the portable AI analytics device is equipped with at least one high-speed Ethernet port (RJ45) for stable, high-speed data communication with the user's laptop; it also features a standard power input interface supporting a wide voltage input to adapt to different power supply conditions. Internally, it integrates an efficient heat dissipation module (such as a fan or heatsink) to ensure that core components remain within a safe operating temperature range under high-load AI computing tasks, guaranteeing system stability and reliability. Essentially, the portable AI analytics device is a highly integrated edge computing node deeply customized and optimized for AI applications. It encapsulates complex AI capabilities in a simple and reliable hardware form, achieving the portability of "computing power."

[0032] In one possible implementation, the computing unit of the portable AI analysis device can be an NVIDIA Jetson development kit, selected from models such as Jetson Nano, Jetson Xavier NX, Jetson Orin Nano, or Jetson Orin NX, depending on performance requirements and cost control. For example, the Jetson Orin NX 16GB version can provide up to 100 TOPS of AI performance, sufficient to smoothly run current mainstream multimodal large models and support parallel analysis of multiple video streams. In terms of integration, the Jetson development kit serves as the motherboard of the entire hardware system, connecting other functional modules through its rich interfaces (such as PCIe, USB, GPIO, etc.). During integration, deep hardware and software adaptation and optimization are performed, including installing a customized Linux operating system (such as the Ubuntu-based JetPack SDK) and configuring NVIDIA's native AI acceleration libraries such as CUDA, cuDNN, and TensorRT, to ensure that multimodal large models can run efficiently and stably on the Jetson platform, maximizing their computing potential.

[0033] The portable AI analysis device is equipped with a network cable interface, a power interface, and a heat dissipation design. It will have at least one Gigabit Ethernet port (RJ45). The choice of wired network connection over Wi-Fi is primarily for data transmission stability and security. A stable, high-speed data link is crucial for ensuring a good user experience when analyzing large amounts of high-definition images and videos. Wired connections also facilitate integration with intranets or dedicated secure networks. For the power interface, a standard DC power input interface will be used, along with a built-in wide-voltage input power management module, making it compatible with 12V or 24V automotive power supplies and general AC adapters, greatly enhancing the device's applicability in different scenarios. Heat dissipation design is paramount to ensuring stable operation over extended periods. Considering the considerable heat generated by the Jetson chip when running AI models at full speed, a combination of active and passive cooling will be employed. The portable AI analysis device's casing will be made of a thermally conductive metal material with large-area heat sink fins. Internally, based on the results of thermal simulation, one or more silent fans will be rationally arranged to form an efficient airflow channel, ensuring that heat can be quickly dissipated and keeping the temperature of the core chip within a safe operating range.

[0034] In one possible implementation, staff connect the portable AI analysis device to their own terminal device (exemplarily, a laptop computer, etc.) via a network cable and start the system. Subsequently, through the user interface on the laptop, they batch import the images and / or videos involved in the case, which need to be extracted from electronic devices such as mobile phones and computers, into the portable AI analysis device.

[0035] Here, the system integrated in the portable AI analysis device can process data in multiple formats. For example, for the images involved in the case, it can process formats such as JPG and PNG; for the videos involved in the case, it can process formats such as MP4 and AVI.

[0036] In one possible implementation, in order to ensure the uniformity and efficiency of subsequent data processing, the received images and / or videos involved in the case can be preprocessed.

[0037] Specifically, the step "preprocessing the image and / or video involved in the case to determine the metadata information corresponding to the image and / or video involved in the case" includes: a1: for the image involved in the case, performing size normalization, format conversion and quality assessment processing on the image involved in the case, and extracting the metadata of the image involved in the case; wherein, the quality assessment is to assess the security and integrity of the image involved in the case.

[0038] a2: For the video in question, perform frame extraction processing on the video in question according to a preset frame rate, divide the video in question into at least one image in question, perform size normalization, format conversion and quality assessment processing on each image in question, and extract the metadata of the image in question.

[0039] In one possible implementation, the received image in question is subjected to size normalization processing, and after size normalization processing, the format is converted to obtain an image format that can be processed subsequently, and then the image in question is subjected to quality assessment.

[0040] Here, quality assessment refers to evaluating the security and integrity of the images involved in the case.

[0041] Specifically, this could involve scanning imported images and / or videos related to the case for viruses and verifying their format to ensure data security and integrity, thereby improving the accuracy of subsequent analysis and report generation.

[0042] In another possible implementation, for the video in question, the video in question is first processed by frame extraction according to a preset frame rate to divide the video in question into at least one image in question, and then each image in question is processed by size normalization, format conversion and quality assessment according to the processing flow for the images in question.

[0043] Here, the preset frame rate can be determined based on the video length of the video in question and the data processing precision. For example, the preset frame rate can be 1 frame per second or a key frame.

[0044] Furthermore, after processing the images and videos involved in the case, it is also necessary to extract the metadata of the images and videos involved in the case to provide auxiliary reference information for subsequent processing.

[0045] Here, the metadata that can be extracted from the images and videos involved in the case includes data such as shooting time, location, and device model.

[0046] Furthermore, after processing the images and / or videos involved in the case, the images and / or videos involved in the case can be input into a pre-trained multimodal large model for processing to obtain labels and description information for the images and / or videos involved in the case.

[0047] S102. Input the processed images and / or videos involved in the case and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the target objects contained in the images and / or videos involved in the case, and extract context and scene information based on the text information in the images and / or videos involved in the case, and output at least one object label and description information for the images and / or videos involved in the case.

[0048] In this application embodiment, a multimodal large model can be used to analyze the content presented in the images and / or videos involved in the case from multiple dimensions. Specifically, at the object level, various common and uncommon items are identified; at the scene level, it is determined whether the location is indoors, outdoors, a hospital, a bank, or a high-risk location; at the text level, Optical Character Recognition (OCR) technology is used to extract and understand all text information in the images, including chat logs, contracts, prescriptions, etc.; at the behavior level, the actions, postures, and interactions of the characters are analyzed; and even at the emotion level, the emotional state of the characters is determined. This information from different dimensions is then fused and correlated to perform cross-modal reasoning. For example, associating "a picture of a pill" with "a chat log about depression" can infer that the person involved may have mental health problems. Through this deep, multi-dimensional, and multimodal intelligent understanding, highly targeted and valuable deep clues can be unearthed.

[0049] Here, the processing flow of the multimodal large model includes multiple steps such as key element detection, text information extraction, dangerous scene identification, scene context analysis, and multimodal association reasoning. The specific processing flow of the multimodal large model will be described below.

[0050] Specifically, the step "inputting the processed images and / or videos involved in the case and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the target objects contained in the images and / or videos involved in the case, and extract context and scene information based on the text information in the images and / or videos involved in the case, and output at least one object label and description information for the images and / or videos involved in the case", includes: b1: inputting the images and / or videos involved in the case into a pre-trained multimodal large model, so that the multimodal large model can visually encode the images and / or videos involved in the case to obtain high-dimensional feature vectors, and decode the high-dimensional feature vectors to determine the target objects contained in the images and / or videos involved in the case.

[0051] b2: Identify the text information contained in the images and / or videos involved in the case, analyze the text information according to criminal investigation terminology, determine at least one professional term and target sensitive words contained in the text information, and generate object tags based on at least one professional term and target sensitive words.

[0052] b3: When the target object is a preset identification object, analyze the object behavior of the target object and the context information of the scene where the target object is located according to the preset dangerous behavior label, and generate object label and scene information.

[0053] b4: Based on the identified target object, object tag, and scene information, output at least one object tag and description information for the image and / or video involved in the case.

[0054] In one possible implementation, the image in question, or the image generated by extracting frames from the video in question, is input into a multimodal large model. The multimodal large model uses a visual encoder to segment the image into multiple small patches and convert them into high-dimensional feature vectors. These feature vectors fuse local details and global contextual information. Subsequently, the decoder receives the feature vectors and a set of learnable query vectors, each query vector responsible for predicting a class of target objects.

[0055] For example, when handling cases of unnatural death, the query vector pays special attention to elements such as "medicine packaging," "hospital wristbands," "medical records," and "high railings." Through a cross-attention mechanism, the multimodal large model can accurately output the category and bounding box coordinates of these involved elements. This allows for precise location of key evidence in the scene, such as a promissory note scattered on a table, a transfer record on a screen, or the edge of a rooftop in the background, providing accurate location information for subsequent investigation of case clues.

[0056] Furthermore, after identifying at least one target object, the text information contained in the images and / or videos involved in the case can be further analyzed.

[0057] Specifically, advanced OCR technologies, such as CRNN-based or Transformer-based models, are employed to accurately recognize various types of text in images, including printed medical records, handwritten IOUs, and screenshots of chat logs on mobile phone screens. When the multimodal model detects a "medical document" area, it extracts information such as the hospital name, department, diagnosis, and drug name; when it detects an "IOU," it extracts key fields such as the borrower, amount, date, and signature. The extracted text is then fed into a specialized NLP model fine-tuned with knowledge from the criminal investigation field. This NLP model has learned a large number of professional terms and target-sensitive words, enabling it to identify key information such as "diagnosis of depression." For example, when a relevant drug name is identified, it is automatically associated with the tag "mental health issues." This domain-specific semantic analysis allows the system to extract case clues from the text of images involved in the case that traditional OCR tools cannot find.

[0058] Furthermore, to ensure the accuracy of the identification process, it is also necessary to associate different dangerous behavior tags with the identification of the target object, and then determine the scene information and object tags.

[0059] For example, when an image contains scene objects such as "rooftop," "rooftop," "riverside," "lakeside," or "overpass," a high-risk assessment mechanism is triggered. Specifically, when identifying a "rooftop," details such as "railing height," "climbing marks," and "leftover items" are detected simultaneously; when identifying a "riverside," elements such as "water depth markings," "slippery surfaces," and "warning signs" are analyzed. A composite risk assessment can also be performed by combining time information (e.g., alone on a rooftop late at night) and the person's state (e.g., dazed expression, holding medication). If an image contains all three elements—"rooftop edge," "medication packaging," and "late at night"—a strong association label of "high-risk fall scene" is generated. Similarly, for a "waterside" scene, if details such as "wading shoe prints" or "leftover personal items" are detected, it is marked as a "high-risk drowning scene." This multi-element composite judgment-based dangerous scene recognition can accurately warn of risks such as falls and drowning in cases.

[0060] In another possible implementation, it is also necessary to analyze the scene based on the context information of the scene information to obtain the analyzed scene information.

[0061] For example, when a "hospital corridor" scene is identified from the images in question, it is combined with surrounding elements such as "clinic signs" and "medical equipment" to confirm that this is a medical environment rather than an ordinary building. For "indoor residential" scenes, details such as "pharmacy bottle placement," "calendar markings," and "notes" are analyzed. For instance, a date circled in red on a calendar with "repayment date" written next to it, combined with an "IOU" on the table and a "screenshot of debt collection text messages," can construct a complete context of "debt pressure." If the "rooftop entrance" in the background of the image in question is identified as open, and the time of the photo is close to the last time the person involved in the case was active, it will be marked as a "key clue for investigation." Similarly, for photos of "water areas," the surrounding environment is analyzed to determine whether it is remote, whether there are "no swimming" signs, and whether personal belongings are found, thereby determining whether it is a "daily recreation" or a "potential suicide location."

[0062] Furthermore, based on the identified target object, object tag, and scene information, at least one object tag and descriptive information can be output for the image and / or video involved in the case.

[0063] Specifically, the descriptive information is generated through the following steps: c1: Based on the identified target object, object label and scene information, clue analysis information is generated according to the preset logical reasoning rules and the knowledge graph in the criminal investigation field.

[0064] c2: Based on the identified target object, object label, and scene information, perform special event analysis and generate special event analysis information.

[0065] c3: Generate the description information based on the clue analysis information and the special event analysis information.

[0066] In one possible implementation, after identifying the target object, object tag, and scene information, clue analysis information can be generated based on preset reasoning rules and a knowledge graph in the criminal investigation field.

[0067] For example, when a medical document (text modality) showing a diagnosis of severe depression is identified from images or videos involved in the case, and a rooftop scene (visual modality) and a draft suicide note (text modality) are identified from other images, the multimodal model combines preset inference rules and a knowledge graph in the field of criminal investigation, utilizing its internal causal reasoning capabilities to associate these clues as "a risk of falling due to mental health problems." By associating "multiple IOU images" (visual + text) and "screenshots of bank collection notices" (text) from the identification results, the conclusion that "the parties involved in the case are facing enormous financial pressure" can be inferred. A time dimension can also be included, allowing the model to arrange these clues chronologically based on the EXIF ​​timestamps extracted from the metadata of the images involved in the case, reconstructing the complete financial chain breakdown process from "normal borrowing" to "falling into usury." This reasoning ability from scattered evidence to a complete case chain greatly enhances the value of clues and the efficiency of case handling.

[0068] In one possible implementation, the multimodal large model can also analyze specific events to enhance the analytical capabilities for criminal cases.

[0069] For example, when the multimodal big data model identifies a "rooftop" or "skyscraper" scene from images and / or videos related to the case, it pays particular attention to details such as whether the railing height is below safety standards, whether there are signs of climbing, and whether personal belongings (such as shoes or suicide notes) have been left behind. For "water" scenes, it analyzes water flow speed, water depth markings, signs of struggle on the shore, and whether valuables such as mobile phones are found. When images of medications such as "sleeping pills" are detected, it combines the quantity of the medication (large-scale stockpiling) and the prescribed dosage to determine whether there is a risk of overdose. In addition, the multimodal big data model can also identify keywords in "search history screenshots," negative comments in "social media screenshots," and mental illness diagnosis history in "medical records." For example, when it identifies that a "photo of the edge of the rooftop" in an image related to the case was taken at 3 a.m., combined with text in a "chat history saying goodbye to a friend" on a mobile phone, a strong warning label for special events is generated. This specific clue mining mechanism for special events can accurately identify risk factors and provide key evidence to support the determination of the nature of the case.

[0070] Furthermore, after identifying at least one object label and descriptive information in the images and / or videos involved in the case through the multimodal large model, at least one object label and descriptive information can be input into the large language model to output the corresponding target report through the large language model.

[0071] S103. Input at least one of the object labels and description information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and description information, and fills the obtained structured data into a report template that matches the selected case type, and outputs a target report corresponding to the case-related image and / or the case-related video.

[0072] In this embodiment, the large language model automatically filters and organizes key information based on the report templates for different case types, combined with at least one object label and descriptive information output by the multimodal large model, to generate a logically clear and formatted analysis report.

[0073] For example, when handling a target case, a target report can be automatically generated, which includes multiple chapters such as "personal health status", "financial status", "social relationships", "recent activity trajectory" and "abnormal behavior records". Each chapter lists relevant picture and video evidence and their analysis summary, which reduces the time that staff spend analyzing and writing reports, thereby comprehensively improving the efficiency and quality of report generation.

[0074] In one possible implementation, since the object labels and description information output by the multimodal large model may be duplicated, and since different object labels have different impacts on subsequent report generation, in order to improve subsequent processing efficiency, the object labels and description information can be deduplicated and filtered before the large language model is processed.

[0075] Specifically, before the step "to make the large language model perform structured processing on the object labels and description information", the generation method further includes: d1: inputting at least one object label and description information into a pre-trained large language model, and filtering and deduplicating at least one object label and description information according to the preset label weight and importance information of each object label to obtain at least one processed object label and description information, so as to perform structured processing on at least one processed object label and description information.

[0076] In this embodiment of the application, different weights can be assigned to different object tags according to different case types, and different object tags and descriptive information can be evaluated according to their importance.

[0077] For example, tags highly relevant to the specific case type being determined (such as "rooftop edge," "deep water area," and "sleeping pill prescription") are assigned the highest weight, while some general tags (such as "indoor," "ground," and "window") may be filtered out. Simultaneously, the large language model leverages its powerful semantic understanding capabilities to identify and merge semantically repetitive or highly similar case-related information. For instance, multiple pieces of evidence, such as "screenshot of hospital diagnosis," "photo of medicine packaging," and "medical appointment text message," are merged into the core clue of "medical records." Through this filtering and deduplication process, the large language model effectively eliminates noise and redundancy, retaining the most crucial and valuable evidence, significantly reducing information overload and laying the foundation for subsequent logical organization.

[0078] Furthermore, after filtering at least one object label and its description, the object label and its description can be structured using a large language model.

[0079] Specifically, the step "inputting at least one of the object labels and description information into a pre-trained large language model so that the large language model performs structured processing on the object labels and description information" includes: e1: inputting at least one of the object labels and description information into a pre-trained large language model so that the large language model determines the hierarchy and association between each of the object labels and description information according to the case logic relationship, and generates structured information based on at least one of the object labels and description information and the hierarchy and association between each of the object labels and description information.

[0080] In one possible implementation, for at least one object label and descriptive information, the large language model can determine the hierarchy and association between the object label and descriptive information according to the case logic relationship, and then generate structured information based on the determined hierarchy and association information.

[0081] For example, the large language model aggregates all evidence related to "medical treatment" (drug packaging, medical records, hospital scene photos) into the "personal health status" dimension; aggregates all evidence related to "economic disputes" (multiple IOUs, collection records, transfer screenshots) into the "financial status and stress sources" dimension; and aggregates all evidence related to "high-risk locations" (rooftop panoramas, rooftop entrances, riverside walkways) into the "risk scenario investigation" dimension. Within each dimension, evidence is further sorted by time or importance; for example, "rooftop photos from the past week" take precedence over "rooftop photos from a month ago," and "a clearly marked psychiatric diagnosis from XX hospital" takes precedence over "a blurry image of a medicine box." Through deep semantic integration of the case, it clearly defines the relationships between various evidence entities involved in the case, providing direct and usable case data raw materials for subsequent report template filling.

[0082] Furthermore, after identifying the structured information, it can be aggregated and categorized, and then the aggregated and categorized information can be filled into the report template to output the target report.

[0083] Specifically, the step "filling the obtained structured data into the report template that matches the selected case type and outputting the target report corresponding to the case-related image and / or the case-related video" includes: f1: aggregating and classifying the structured data according to the analysis framework and evidence classification rules corresponding to the case type, and determining the classified type structured data.

[0084] f2: Fill the target area of ​​the report template with the determined type of structured data and the corresponding case-related images and / or videos according to the preset display format to obtain the initial report.

[0085] f3: Generate a case information summary based on the data filled in each target area in the initial report.

[0086] f4: Based on the initial report and the case information summary, output a target report corresponding to the case-related image and / or the case-related video.

[0087] In this application embodiment, different analysis frameworks and evidence classification rules are adopted for different case types. The structured data can be classified according to the analysis framework and evidence classification rules corresponding to the case type to determine the classified type of structured data.

[0088] For example, evidence related to "medical records" (screenshots of psychiatric diagnoses, psychological counseling appointment records) is precisely categorized under the "Personal Health Status" section; "rooftop photos" (measurements of railing height, climbing marks, left-behind shoes or suicide notes) and "wading photos" (water depth markers, personal belongings on the shore, slippery steps) are categorized under the "High-Risk Location Investigation" section. Furthermore, the large language model performs deep aggregation based on the spatiotemporal relevance of this evidence: for instance, it aggregates the time the "rooftop photo" was taken, the "search history" on the phone, and the "medical records"—three cross-modal pieces of evidence—into the high-value clue of "suspected premeditated fall." Similarly, it aggregates "riverside photos," "phone browsing history," and "chat records" into "suspected premeditated drowning." This intelligent aggregation and categorization of evidence based on case type ensures that the final report accurately presents the core elements of the case, forming a complete chain of evidence and providing strong support for determining the nature of the case.

[0089] Furthermore, the obtained structured data and corresponding images and / or videos involved in the case can be filled into the target area of ​​the report template according to a preset display format to obtain an initial report.

[0090] In one possible implementation, the report generation process may include: selection and matching of report templates, intelligent filling of structured data, and extraction of final summary and conclusions.

[0091] Firstly, regarding the selection and matching of report templates: a rich library of report templates can be pre-set, covering various types of cases commonly encountered in criminal investigation work. Each report template has a rigorous structure and comprehensive content. When the user selects a case type during the task configuration phase, the corresponding report template is automatically matched. For example, if the user selects "unnatural death cases," the system will load a report template containing sections such as "Personal Basic Information," "Health Status Analysis," "Financial Status Analysis," "Social Relationship Analysis," "Abnormal Behavior Records," and "Scene Environment Analysis."

[0092] Secondly, regarding intelligent data filling: After the report template is determined, the large language model will fill the corresponding parts of the report template with the generated structured data. The large language model will transform the structured data points (such as tags, entities, and relationships) into fluent, professional, and document-compliant natural language descriptions according to the specific requirements of each chapter and key point in the report template.

[0093] For example, for structured data "{object: medicine bottle, text: "XX brand", associated person: party involved}", the large language model might generate the following description: "A bottle of medicine named 'XX brand' was found at the scene. Upon investigation, this medicine is a selective serotonin reuptake inhibitor, commonly used for specific mental illnesses. Combined with the negative comments posted by the party involved on social media accounts, it is preliminarily determined that the party involved may have mental health issues, and further investigation of their medical records is required." Simultaneously with text generation, relevant images or videos related to the case are automatically inserted into the report in a preset display format, such as thumbnails, forming an initial report with both text and images.

[0094] Thirdly, regarding the extraction of the final summary and conclusions: In one possible implementation, the report also needs to include a highly summarized case summary and conclusive opinions. The case information summary can be generated by combining the data filled in each target area of ​​the entire initial report, thereby obtaining the final target report.

[0095] For example, at the end of a report on an unnatural death, the large language model might write: "Based on a comprehensive analysis of the existing evidence, the person faced immense financial pressure and had severe depressive tendencies before death. The location of death was an abandoned construction site, unattended, and no signs of struggle were found at the scene. Combined with the photo of the suicide note found on their phone, it is preliminarily determined that the person is most likely XX. The recommended next steps are: 1. Verify the authenticity of the suicide note; 2. Investigate their debtors; 3. Review surveillance footage from the area surrounding the abandoned construction site to confirm their last known activity."

[0096] Furthermore, the generated target reports, which are structurally complete, detailed in content, and logically clear, can be output in standard formats such as PDF or Word, making them easy to archive, print, and share.

[0097] In one possible implementation, large multimodal models and large language models, along with their runtime environments (including operating systems, dependency libraries, etc.), will be encapsulated in container images such as Docker. This containerized deployment approach offers numerous advantages: it ensures the consistency and reproducibility of the software environment; it simplifies model installation, startup, and shutdown with just a few commands; and it facilitates model version management and rollback. 2. Model Compression and Optimization: To efficiently run large AI models on resource-constrained Jetson edge devices, a series of model compression and optimization techniques are employed. For example, using TensorRT to quantize the model with INT8 or FP16 precision can significantly improve inference speed and reduce memory usage with almost no loss of accuracy. Furthermore, knowledge distillation and pruning techniques are used to further reduce model size while maintaining performance. 3. Online / Offline Update Mechanism: The system supports online and offline model updates. When a new version of the model or a new analysis template is released, the update package can be pushed to portable AI analysis devices via the intranet or a dedicated USB flash drive. The system automatically deploys and switches to the new model, and the entire process is transparent to the user. This flexible update mechanism ensures that the technical solutions of the embodiments of this application can continuously adapt to new case analysis needs.

[0098] In another possible implementation, while keeping the overall technical concept of the present application unchanged, an equivalent alternative can also implement the method for generating the analysis report in the embodiments of the present application.

[0099] Specifically, existing edge computing platforms containing different platforms or chips can be used to replace portable AI analysis devices; or the functions of portable AI analysis devices can be directly integrated into devices in the field of criminal investigation. This can be done by integrating a dedicated AI coprocessor (such as an NPU) on the motherboard of a tablet computer, or by using the computing power of its main CPU / GPU to run lightweight AI models.

[0100] In another possible implementation, the multimodal model processing can be performed outside the portable AI analysis device. Instead, the portable AI analysis device acts as a front-end for data preprocessing and result display, while the core AI inference tasks are handled by a large multimodal model deployed in the cloud or intranet data center. In this mode, the portable AI analysis device performs preliminary preprocessing such as compression and feature extraction on the images and videos involved in the case. Then, it securely sends the processed data to a cloud server via an encrypted network connection (such as a VPN). The cloud server runs a more powerful and up-to-date large multimodal model, which, after completing complex analysis tasks, returns the results to the portable AI analysis device at the front end.

[0101] The specific process is as follows: Portable AI analysis devices at the front end are only responsible for performing preliminary analysis tasks with relatively low computational load but high real-time requirements. For example, running a lightweight multimodal model to quickly extract key features, tags, and text information from the images and videos involved in the case. Then, these structured preliminary analysis results are sent via the network to a large language model deployed on a cloud or intranet server. The server-side large language model, with its more powerful computing capabilities and richer knowledge base, is responsible for in-depth integration, logical reasoning, and report writing of the massive amounts of preliminary analysis results. This "edge computing + cloud intelligence" model can be seen as a compromise between the aforementioned "all-cloud" and "all-edge" solutions. It reduces the computational burden on front-end devices and lowers the dependence on network bandwidth (because structured results are transmitted instead of raw images and videos), while utilizing the powerful large language model capabilities of the cloud to ensure report quality. This alternative workflow achieves a better balance between system performance and resource consumption, making it a more pragmatic and efficient choice. To further improve the accuracy and reliability of the analysis results and increase staff participation and trust in the results, a key human-computer interaction step can be added to the workflow. Before the portable AI analysis device completes its initial analysis but generates the final report, the system can display all identified tags, keyframes, text fragments, and other intermediate results to staff through an interactive interface. Staff can perform the following operations on this interface: 1. Filtering and Confirmation: Manually review the results identified by the AI, removing obviously erroneous tags and confirming important clues. 2. Supplementation and Correction: Staff can manually add or modify information that the AI ​​failed to identify or misidentified. 3. Highlighting: Highlight or annotate particularly important clues. After manual review, the large language model generates the final report based on the results verified and supplemented by the "human-machine hybrid" process. The advantages of this process are: 1. Improved Accuracy: Minimizes misjudgments and omissions by the AI ​​model. 2. Enhanced Interpretability: Staff can clearly see each step of the AI ​​analysis process, understand the origin of the report's conclusions, and thus increase their trust in the system.

[0102] In another possible implementation, for multimodal large models and large language models, LLaVA, MiniGPT-4, BLIP-2, OpenAI's GPT-4V, and Google's Gemini Pro Vision can all be used as processing models in this application. For example, an open-source multimodal model can be used for basic visual understanding, and its output can then be fed into a powerful large language model (such as GPT-4) for final report polishing and generation. Alternatively, to achieve complete autonomy and control, all open-source models that have been fine-tuned for the specific domain can be used. This flexibility in model combination allows the technical solution of this invention to adapt to different application scenarios and constraints, exhibiting strong scalability and vitality.

[0103] The method for generating an analysis report provided in this application includes receiving extracted images and / or videos related to a case, preprocessing the images and / or videos to determine the metadata information corresponding to them, inputting the processed images and / or videos and metadata information into a pre-trained multimodal large model to enable the model to identify target objects contained in the images and / or videos, extracting context and scene information based on text information in the images and / or videos, and outputting at least one object label and description information for the images and / or videos. The method also includes inputting the at least one object label and description information into a pre-trained large language model to enable the model to perform structured processing on the object labels and description information, filling the obtained structured data into a report template matching the selected case type, and outputting a target report corresponding to the images and / or videos. In this way, by using a portable AI analysis device to input the extracted images and / or videos related to the case into a multimodal big model, the images and / or videos related to the case are analyzed from multiple dimensions to obtain at least one object label and descriptive information. Then, the at least one object label and descriptive information are input into a big language model, and the corresponding target report is directly generated through the processing in the big language model, which can improve the comprehensiveness and efficiency of report generation.

[0104] Based on the same inventive concept, this application also provides an analysis report generation device corresponding to the analysis report generation method. Since the principle of the device in this application is similar to the analysis report generation method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0105] Please refer to Figures 2 and 3. Figure 2 is a schematic diagram of one structure of an analysis report generation device provided in an embodiment of this application, and Figure 3 is a schematic diagram of another structure of an analysis report generation device provided in an embodiment of this application. As shown in Figure 2, the generation device 200 includes: a data preprocessing module 210, used to receive extracted images and / or videos related to the case, preprocess the images and / or videos related to the case, and determine the metadata information corresponding to the images and / or videos related to the case; and a tag and description information extraction module 220, used to input the processed images and / or videos related to the case and the metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the target objects contained in the images and / or videos related to the case, and based on the case... The text information in the image and / or the video involved in the case is used to extract context and scene information, and output at least one object label and descriptive information for the image and / or the video involved in the case; the target report generation module 230 is used to input at least one of the object labels and descriptive information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and descriptive information, and fills the obtained structured data into a report template that matches the selected case type, and outputs a target report corresponding to the image and / or the video involved in the case.

[0106] In one possible implementation, when the tag and description information extraction module 220 is used to input the processed image and / or video of the case and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify target objects contained in the image and / or video of the case, and extract context and scene information based on the text information in the image and / or video of the case, and output at least one object tag and description information for the image and / or video of the case, the tag and description information extraction module 220 is configured to: input the image and / or video of the case into the pre-trained multimodal large model, so that the multimodal large model can perform visual encoding on the image and / or video of the case to obtain a high-dimensional feature vector. The high-dimensional feature vector is decoded to determine the target object contained in the image and / or video involved in the case; text information contained in the image and / or video involved in the case is identified, and the text information is analyzed according to criminal investigation terminology to determine at least one professional term and target sensitive words contained in the text information, and object tags are generated based on at least one professional term and target sensitive words; when the target object is a preset identification object, the object behavior of the target object and the context information of the scene where the target object is located are analyzed according to preset dangerous behavior tags to generate object tags and scene information; based on the identified target object, object tags and scene information, at least one object tag and descriptive information for the image and / or video involved in the case are output.

[0107] In one possible implementation, the tag and description information extraction module 220 is used to generate description information through the following steps: generating clue analysis information based on the identified target object, object tag, and scene information, according to preset logical reasoning rules and a knowledge graph in the criminal investigation field; performing special event analysis based on the identified target object, object tag, and scene information to generate special event analysis information; and generating the description information based on the clue analysis information and the special event analysis information.

[0108] In one possible implementation, as shown in FIG3, the generation device 200 further includes a tag filtering module 240, which is used to: input at least one object tag and description information into a pre-trained large language model, and filter and deduplicate at least one object tag and description information according to the preset tag weight and importance information of each object tag to obtain at least one processed object tag and description information, so as to perform structured processing on at least one processed object tag and description information.

[0109] In one possible implementation, when the target report generation module 230 is used to input at least one of the object tags and description information into a pre-trained large language model so that the large language model performs structured processing on the object tags and description information, the target report generation module 230 is used to: input at least one of the object tags and description information into a pre-trained large language model so that the large language model determines the hierarchy and association between each of the object tags and description information according to the case logic relationship, and generates structured information based on at least one of the object tags and description information and the hierarchy and association between each of the object tags and description information.

[0110] In one possible implementation, when the target report generation module 230 is used to fill the obtained structured data into a report template matching the selected case type and output a target report corresponding to the case-related image and / or the case-related video, the target report generation module 230 is configured to: aggregate and classify the structured data according to the analysis framework and evidence classification rules corresponding to the case type, and determine the classified type structured data; fill the determined type structured data and the corresponding case-related image and / or case-related video into the target area of ​​the report template according to a preset display format to obtain an initial report; generate a case information summary based on the data filled in each target area of ​​the initial report; and output a target report corresponding to the case-related image and / or the case-related video based on the initial report and the case information summary.

[0111] In one possible implementation, when the data preprocessing module 210 preprocesses the image and / or video involved in the case to determine the metadata information corresponding to the image and / or video involved in the case, the data preprocessing module 210 is configured to: for the image involved in the case, perform size normalization, format conversion, and quality assessment processing on the image involved in the case, and extract the metadata of the image involved in the case; wherein, the quality assessment is to assess the security and integrity of the image involved in the case; for the video involved in the case, perform frame extraction processing on the video involved in the case according to a preset frame rate, divide the video involved in the case into at least one image involved in the case, perform size normalization, format conversion, and quality assessment processing on each image involved in the case, and extract the metadata of the image involved in the case.

[0112] The analysis report generation apparatus provided in this application embodiment receives extracted case-related images and / or case-related videos, preprocesses the case-related images and / or case-related videos to determine the metadata information corresponding to the case-related images and / or case-related videos; inputs the processed case-related images and / or case-related videos and metadata information into a pre-trained multimodal large model, so that the multimodal large model can identify the target objects contained in the case-related images and / or case-related videos, and extract context and scene information based on the text information in the case-related images and / or case-related videos, and outputs at least one object label and descriptive information for the case-related images and / or case-related videos; inputs at least one object label and descriptive information into a pre-trained large language model, so that the large language model can perform structured processing on the object label and descriptive information, and fills the obtained structured data into a report template matching the selected case type, and outputs a target report corresponding to the case-related images and / or case-related videos. In this way, by using a portable AI analysis device to input the extracted images and / or videos related to the case into a multimodal big model, the images and / or videos related to the case are analyzed from multiple dimensions to obtain at least one object label and descriptive information. Then, the at least one object label and descriptive information are input into a big language model, and the corresponding target report is directly generated through the processing in the big language model, which can improve the comprehensiveness and efficiency of report generation.

[0113] Please refer to Figure 4, which is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. As shown in Figure 4, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.

[0114] The memory 420 stores machine-readable instructions that can be executed by the processor 410. When the electronic device 400 is running, the processor 410 and the memory 420 communicate via the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the analysis report generation method in the method embodiment shown in Figure 1 above can be executed. For specific implementation, please refer to the method embodiment, which will not be repeated here.

[0115] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the analysis report generation method in the method embodiment shown in FIG1 above. For specific implementation details, please refer to the method embodiment, which will not be repeated here.

[0116] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0117] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0118] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0119] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0120] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating an analysis report, characterized in that, The generation method, applicable to portable AI analysis devices, includes: receiving extracted images and / or videos related to a case; preprocessing the images and / or videos to determine metadata information corresponding to the images and / or videos; inputting the processed images and / or videos and metadata information into a pre-trained multimodal large model, enabling the multimodal large model to identify target objects contained in the images and / or videos, extract context and scene information based on text information in the images and / or videos, and output at least one object label and description information for the images and / or videos; inputting at least one object label and description information into a pre-trained large language model, enabling the large language model to perform structured processing on the object labels and description information, and filling the obtained structured data into a report template matching the selected case type, and outputting a target report corresponding to the images and / or videos.

2. The generation method according to claim 1, characterized in that, The step of inputting the processed images and / or videos involved in the case, along with metadata information, into a pre-trained multimodal large model, enabling the multimodal large model to identify target objects contained in the images and / or videos involved in the case, and to extract context and scene information based on text information in the images and / or videos involved in the case, outputting at least one object label and descriptive information for the images and / or videos involved in the case, includes: inputting the images and / or videos involved in the case into the pre-trained multimodal large model, enabling the multimodal large model to visually encode the images and / or videos involved in the case to obtain high-dimensional feature vectors, and decoding the high-dimensional feature vectors to determine the images involved in the case. The system identifies target objects contained in the images and / or videos involved in the case; it identifies text information contained in the images and / or videos involved in the case, analyzes the text information according to criminal investigation terminology, determines at least one professional term and target sensitive words contained in the text information, and generates object tags based on at least one professional term and target sensitive words; when the target object is a preset identification object, it analyzes the object behavior of the target object and the context information of the scene where the target object is located according to preset dangerous behavior tags, and generates object tags and scene information; based on the identified target object, object tags and scene information, it outputs at least one object tag and descriptive information for the images and / or videos involved in the case.

3. The generation method according to claim 2, characterized in that, The descriptive information is generated through the following steps: Based on the identified target object, object tags, and scene information, clue analysis information is generated according to preset logical reasoning rules and knowledge graphs in the field of criminal investigation; Based on the identified target objects, object tags, and scene information, special event analysis is performed to generate special event analysis information; The description information is generated based on the clue analysis information and the special event analysis information.

4. The generation method according to claim 1, characterized in that, Before the large language model performs structured processing on the object labels and description information, the generation method further includes: inputting at least one object label and description information into a pre-trained large language model, and filtering and deduplicating at least one object label and description information according to the preset label weight and importance information of each object label to obtain at least one processed object label and description information, so as to perform structured processing on at least one processed object label and description information.

5. The generation method according to claim 1, characterized in that, The step of inputting at least one of the object labels and description information into a pre-trained large language model, so that the large language model performs structured processing on the object labels and description information, includes: inputting at least one of the object labels and description information into a pre-trained large language model, so that the large language model determines the hierarchy and association between each of the object labels and description information according to the logical relationship of the case, and generates structured information based on at least one of the object labels and description information and the hierarchy and association between each of the object labels and description information.

6. The generation method according to claim 1, characterized in that, The step of filling the obtained structured data into a report template matching the selected case type and outputting a target report corresponding to the case-related image and / or video includes: aggregating and classifying the structured data according to the analysis framework and evidence classification rules corresponding to the case type to determine the classified type structured data; filling the determined type structured data and the corresponding case-related image and / or video into the target area of ​​the report template according to a preset display format to obtain an initial report; generating a case information summary based on the data filled in each target area of ​​the initial report; and outputting a target report corresponding to the case-related image and / or video based on the initial report and the case information summary.

7. The generation method according to claim 1, characterized in that, The preprocessing of the images and / or videos involved in the case to determine the corresponding metadata information includes: for the images involved in the case, performing size normalization, format conversion, and quality assessment on the images involved in the case, and extracting the metadata of the images involved in the case; wherein, the quality assessment is to evaluate the security and integrity of the images involved in the case; for the videos involved in the case, performing frame extraction processing on the videos involved in the case according to a preset frame rate, dividing the videos involved in the case into at least one image involved in the case, performing size normalization, format conversion, and quality assessment on each image involved in the case, and extracting the metadata of the images involved in the case.

8. An apparatus for generating an analysis report, characterized in that, The generation device, applied to portable AI analysis equipment, comprises: a data preprocessing module for receiving extracted images and / or videos related to a case, preprocessing the images and / or videos, and determining the metadata information corresponding to the images and / or videos; a tag and description information extraction module for inputting the processed images and / or videos and metadata information into a pre-trained multimodal large model, enabling the multimodal large model to identify target objects contained in the images and / or videos, extract context and scene information based on text information in the images and / or videos, and output at least one object tag and description information for the images and / or videos; and a target report generation module for inputting at least one object tag and description information into a pre-trained large language model, enabling the large language model to perform structured processing on the object tags and description information, and filling the obtained structured data into a report template matching the selected case type, and outputting a target report corresponding to the images and / or videos.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the method for generating an analysis report as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for generating an analysis report according to any one of claims 1 to 7.