Multi-modal research data comprehensive analysis method, system and equipment and storage medium

Through the multimodal identification model and predefined calculation rules, the problem of multimodal data integration is solved, and the comprehensive analysis of multimodal research data is realized, which improves the analysis efficiency and reliability.

CN120492886APending Publication Date: 2025-08-15INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524828.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional analysis methods are difficult to effectively integrate multimodal data, resulting in the lack of a cross-modal evaluation index system, affecting the efficiency and reliability of in-depth analysis of research data.

Method used

A multimodal recognition model is used to identify target type parameters from multiple modal data samples, and a target type index value is generated through predefined calculation rules to realize a comprehensive analysis of multimodal research data.

Benefits of technology

It realizes effective information extraction and unified parameter evaluation of multimodal research data, and improves the analysis efficiency and reliability of multimodal research data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492886A_ABST
    Figure CN120492886A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and particularly provides a multi-modal survey data comprehensive analysis method, system and device and a storage medium, and the method comprises the steps: obtaining a plurality of data samples, the data samples comprising data of multiple modals; identifying a target type parameter from the data sample by using a multi-modal identification model, and counting a parameter value of the target type parameter in the data sample; and generating a target type index value of each data sample in the overall data according to the parameter values of the plurality of data samples and a predefined calculation rule. According to the method, comprehensive analysis of the multi-modal investigation data is realized, and the analysis efficiency of the multi-modal investigation data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data analysis technology, and specifically relates to a method, system, device and storage medium for comprehensive analysis of multimodal survey data. Background Art

[0002] In the field of survey data analysis, the key technical challenge currently faced is that when the survey objects include multi-modal data such as videos, images, and text, traditional analysis methods find it difficult to effectively integrate multi-source heterogeneous information.

[0003] Due to the lack of unified quantitative standards for data of different modalities, it is impossible to establish a cross-modal evaluation index system during the comprehensive evaluation process. Specifically, the semantic gap between multimodal data is difficult to bridge, and the characteristic parameters of each modality cannot be mapped to a unified calculation framework, which makes it difficult to automatically generate target type indicator values for data samples in the overall analysis.

[0004] The lack of multimodal data fusion and indicator quantification capabilities directly restricts the efficiency and reliability of in-depth analysis of survey data. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method, system, device and storage medium for comprehensive analysis of multimodal survey data to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a method for comprehensive analysis of multimodal survey data, comprising: Acquire multiple data samples, where the data samples include data of multiple modalities; Using a multimodal recognition model to identify target type parameters from data samples, and counting parameter values of the target type parameters in the statistical data samples; According to the parameter values of the plurality of data samples and predefined calculation rules, a target type index value of each data sample in the overall data is generated.

[0007] In an optional embodiment, the method further comprises: Collecting a first data sample of a first object at a first time, and calculating a first target type indicator value based on the first data sample; collecting a second data sample of the first object at a second time, and calculating a second target type indicator value based on the second data sample; Calculating a difference between the second target type index value and the first target type index value, and generating evaluation information of the first object based on the difference; The second time is after the first time.

[0008] In an optional embodiment, a plurality of data samples are obtained, wherein the data samples include data of multiple modalities, including: A plurality of data samples corresponding to a plurality of objects are obtained, where the data samples include video data, text data, and image data.

[0009] In an optional embodiment, the method further comprises: The data samples of multiple objects corresponding to the same parent object are divided into the same group, and the data samples in the same group are analyzed.

[0010] In an optional embodiment, a multimodal recognition model is used to identify target type parameters from data samples, and parameter values of the target type parameters in the data samples are counted, including: Store the target type parameters identified by the multimodal recognition model from the data sample in a specified format; The stored target type parameters are deduplicated, and the parameter values of the deduplicated target type parameters are counted.

[0011] In an optional embodiment, generating a target type index value for each data sample based on the parameter values of the plurality of data samples and a predefined calculation rule includes: Predefined calculation rules:

[0012] Where y is the target type indicator value of the object to which the data sample belongs; x is the parameter value of the target type parameter of the data sample; N is the variation factor, including the integer part and the decimal part. The integer part is the number of digits in the sum of the parameter values of the objects in the same group, and the decimal part is the first digit of the sum of the parameter values of the objects in the same group. The parameter values of the target type parameters of the multiple objects in the same group are substituted into the predefined calculation rule to obtain the target type index value of each object.

[0013] In an optional embodiment, the method further comprises: If it is determined that each object includes multiple target types, the target type index value corresponding to each target type is calculated respectively; Calculate the weighted sum of multiple target type index values of the same object to obtain the total index value of the object within the group.

[0014] In a second aspect, the present invention provides a multimodal survey data comprehensive analysis system, comprising: An acquisition module, configured to acquire a plurality of data samples, wherein the data samples include data of multiple modalities; An identification module is used to identify target type parameters from data samples using a multimodal recognition model and to count parameter values of the target type parameters in the data samples; The analysis module is used to generate a target type index value for each data sample in the overall data based on the parameter values of the multiple data samples and predefined calculation rules.

[0015] According to a third aspect, a device is provided, comprising: Memory, used to store a comprehensive analysis program for multimodal survey data; A processor is configured to implement the steps of the multimodal survey data comprehensive analysis method provided in the first aspect when executing the multimodal survey data comprehensive analysis program.

[0016] In a fourth aspect, a computer-readable storage medium is provided, on which a multimodal survey data comprehensive analysis program is stored. When the multimodal survey data comprehensive analysis program is executed by a processor, the steps of the multimodal survey data comprehensive analysis method provided in the first aspect are implemented.

[0017] The beneficial effects of the present invention lie in the fact that the method, system, device, and storage medium for comprehensive analysis of multimodal survey data provided herein extract effective information from multimodal survey data through a multimodal recognition model, unify the effective parameters of survey data from different modalities, and then evaluate the overall ranking of different samples based on predefined calculation rules and the effective parameters of different samples. This invention enables comprehensive analysis of multimodal survey data and improves the efficiency of multimodal survey data analysis.

[0018] In addition, the present invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0021] Figure 2 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.

[0022] Figure 3 A schematic structural diagram of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0025] The multimodal survey data comprehensive analysis method provided in the embodiment of the present invention is executed by a computer device. Accordingly, the multimodal survey data comprehensive analysis system runs in the computer device.

[0026] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution entity can be a multimodal survey data comprehensive analysis system. According to different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0027] like Figure 1 As shown, the method includes: S1. Acquire multiple data samples, where the data samples include data of multiple modalities.

[0028] Data modality definition: Identify different data modalities, such as images, text, audio, and video. Each modality has different characteristics and formats. For example, image data typically exists as a matrix of pixels, text data is a sequence of characters, audio data is a time-series sound signal, and video data is a continuous sequence of image frames.

[0029] Data collection method: Image data: can be obtained from public image datasets, collected in real time by cameras, or crawled from social media platforms, image sharing websites, etc.

[0030] Text data: It can be captured from news websites, blogs, forums, etc., or public text corpora can be used.

[0031] Audio data: can be downloaded from an audio library or recorded via a microphone.

[0032] Video data: can be downloaded from video websites or recorded using surveillance cameras.

[0033] Data sample organization: Assign a unique identifier to each data sample to associate data from different modalities. For example, you can use a dictionary to store each data sample, with the key being the modality name and the value being the corresponding data file path or data object.

[0034] Perform preliminary cleaning and preprocessing on the data, such as removing noise in images, special characters in text, background noise in audio, etc.

[0035] S2. Use the multimodal recognition model to identify target type parameters from the data samples, and count the parameter values of the target type parameters in the statistical samples.

[0036] Multimodal recognition model selection: Fusion model architecture: You can choose a multimodal fusion model based on deep learning, such as early fusion, late fusion, or hybrid fusion. Early fusion concatenates data from different modalities at the input layer, while late fusion fuses features from different modalities at the output layer. Hybrid fusion combines the advantages of both early and late fusion.

[0037] Pre-trained models: Using pre-trained multimodal models (such as CLIP and Multimodal BERT) can reduce the time and cost of model training. These pre-trained models are trained on large-scale multimodal data and have strong feature extraction capabilities.

[0038] Model training and fine-tuning: If you use a pre-trained model, you need to fine-tune it on your own dataset. First, divide the data samples into training, validation, and test sets. Then, fine-tune the model on the training set, using the backpropagation algorithm to update the model parameters to minimize the loss function.

[0039] During training, data augmentation techniques can be used to increase data diversity and improve the generalization ability of the model. For example, operations such as rotating, flipping, and scaling images, and synonym replacement, insertion, and deletion can be performed on text.

[0040] Target type parameter identification: Define target type parameters. For example, in image-text multimodal data, the target type parameter can be the object category in the image, the emotional tendency in the text, etc.

[0041] The data samples are input into the trained multimodal recognition model, and the model outputs the predicted results of the target type parameters for each data sample.

[0042] Parameter value statistics: For each data sample, calculate the target type parameter value based on the model's output. For example, if the target type parameter is object category, you can count the number of occurrences of each category; if the target type parameter is sentiment, you can count the proportion of positive, negative, and neutral sentiment.

[0043] S3. Generate a target type index value for each data sample in the overall data based on the parameter values of the multiple data samples and predefined calculation rules.

[0044] Predefined calculation rules: Develop predefined calculation rules based on specific application scenarios and requirements. For example, you can use weighted average, normalization, standardization, and other methods to calculate the target type indicator value.

[0045] If data from different modalities contribute differently to the target type indicator value, different weights can be assigned to the data from each modality. For example, in image-text multimodal data, if image data contributes more to the target type indicator value, a higher weight can be assigned to the image data.

[0046] Target type indicator value calculation: For each data sample, calculate its target type indicator value in the overall data according to the predefined calculation rules and the parameter values obtained by statistics. For example, if the predefined calculation rule is weighted average, then the parameter value of each target type parameter can be multiplied by the corresponding weight and then summed to obtain the target type indicator value.

[0047] The calculated target type indicator values are post-processed, such as normalizing them to the interval [0, 1], to facilitate comparison and analysis.

[0048] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0049] Data samples of multiple objects corresponding to the same parent object are divided into the same group, and the data samples in the same group are analyzed. A plurality of data samples corresponding to the multiple objects are obtained, wherein the data samples include video data, text data, and image data.

[0050] For example, if the object is a city-level jurisdiction and the parent object is a provincial-level jurisdiction, the data samples of all city-level jurisdictions in the same provincial-level jurisdiction are grouped together. The data sample is the number of enterprises.

[0051] Data source: Video data: This can be obtained from government-issued propaganda videos, surveillance videos, news reports, and other channels, such as official provincial or municipal websites and news media platforms.

[0052] Text data: This includes government-issued statistical reports, policy documents, press releases, and discussions about the jurisdiction on social media. Web crawlers can be used to extract data from relevant websites, forums, and social media platforms.

[0053] Image data: such as satellite maps, cityscape photos, aerial photos of business distribution, etc. This can be obtained from geographic information system (GIS) data providers, official image libraries, or field photography.

[0054] Enterprise quantity data: can be obtained from official channels such as the enterprise registration database of the industrial and commercial administration department and statistical yearbooks.

[0055] Data storage and management: Use a database (such as MySQL, MongoDB, etc.) to store the acquired data. Create a record for each object (such as a city district), including information such as the storage path of the video data, the text data content, the storage path of the image data, and the number of enterprises.

[0056] Add a unique identifier to each data sample to facilitate subsequent query and management.

[0057] The relationship between the parent object and the object: Establish mapping relationships: Create a mapping table to record the correspondence between objects (such as municipal districts) and their parent objects (such as provincial districts). This relationship can be stored in the database for fast query.

[0058] Data reading and grouping: Read a sample of all objects from the database.

[0059] According to the mapping table, data samples of multiple objects (such as municipal districts) corresponding to the same parent object (such as the same provincial district) are divided into the same group.

[0060] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0061] S201. Store the target type parameters identified from the data sample by the multimodal recognition model in a specified format.

[0062] Multimodal recognition models include: Input processing: Receives an image and a natural language instruction (such as detection, recognition, spotting, or semantic understanding).

[0063] Feature Extraction: CLIP-ViT-L / 14 is used as a visual encoder to extract visual features from the image and incorporate the Transformer front and back layer grid features. The extracted feature maps are then flattened into a sequence of visual embeddings and projected into the embedding dimension of the LLM using a linear layer.

[0064] ‌Multimodal Fusion‌: Concatenate the processed visual embedding sequence with the embedding sequence tokenized from language instructions and feed into a large language model such as Vicuna.

[0065] Task execution: Based on the knowledge base of a large language model and combined with the instruction content, the corresponding multimodal understanding task is completed.

[0066] Among them, CLIP-ViT-L / 14 is a multimodal pre-training model that adopts the Vision Transformer (ViT) architecture and has powerful image encoding capabilities.

[0067] The training process is divided into two stages, both of which adopt a unified multi-modal instruction fine-tuning method: ‌Pre-training phase‌: Purpose: Align the output features of a pre-trained visual encoder with the feature space of a large language model.

[0068] ‌Operation‌: Freeze the pre-trained large vision and language models and only train the linear projection layer to align visual and language features.

[0069] Data: The pre-training data consists of 595K filtered natural scene images and their captions from a dataset, and 600K image-text pairs from PowerPoint presentations. We use an OCR tool to accurately extract text and box annotations from each image and construct OCR instructions based on them.

[0070] Fine-tuning phase: Purpose: Further optimize the weights of large language models and add additional multimodal understanding tasks.

[0071] ‌Operation‌: Unfreeze the large language model and projection layers, and further add an additional multimodal understanding task for text-rich images in addition to the training tasks in the pre-training phase.

[0072] ‌Data‌: Based on pre-training data, collect image-text pairs to enhance the model's ability to understand diverse instructions.

[0073] Using the trained multimodal recognition model, we identify the target type parameters of each piece of data from the data sample. For example, we extract company information from a piece of text in the data sample and store it in the company name-address format. This way, all company information for the same data sample is stored in the same list.

[0074] S202. De-duplicate the stored target type parameters, and count the parameter values of the de-duplicated target type parameters.

[0075] For example, by deduplicating the enterprise information in the list of the same data sample and then counting the number of pieces of deduplicated enterprise information, the number of enterprises in the jurisdiction to which the data sample belongs can be obtained.

[0076] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0077] Predefined calculation rules:

[0078] Where y is the target type indicator value of the object to which the data sample belongs; x is the parameter value of the target type parameter of the data sample; N is the variation factor, including the integer part and the decimal part. The integer part is the number of digits in the sum of the parameter values of the objects in the same group, and the decimal part is the first digit of the sum of the parameter values of the objects in the same group. The parameter values of the target type parameters of the multiple objects in the same group are substituted into the predefined calculation rule to obtain the target type index value of each object.

[0079] Take enterprise data analysis as an example: Data definition: x is the "number of enterprises" in each city, which must be a positive integer.

[0080] The total data volume of jurisdiction A (province): 12,345,678 companies.

[0081] Calculation of the variation factor N: Integer part: The number of digits in the integer part of the total data volume of jurisdiction A.

[0082] The total data volume is 12,345,678, and the integer part is 8 bits (1, 2, 3, 4, 5, 6, 7, 8).

[0083] Therefore, the integer part of N is 8.

[0084] Decimal part: the first digit of the integer part of the total data amount.

[0085] The first digit of the total data volume is 1 (i.e., "1" in 12,345,678).

[0086] Therefore, the fractional part of N is 0.1.

[0087] Final N: N=8.1.

[0088] Comprehensive ratio calculation (taking city A as an example): Assume that the data volume x=5,000 (the number of enterprises in the city is 5,000).

[0089] Substituting into the formula:

[0090] The result y=0.447, satisfying y<1.

[0091] Convert to thousandths: 0.447×1000=447 points.

[0092] Horizontal comparison: Assume that the scores of other cities are: City b: y = 0.520 × 1000 = 520 points City C: y = 0.380 × 1000 = 380 points City d: y=0.600×1000=600 points Conclusion: City D performed the best (600 points), while City A (447 points) needs to improve.

[0093] Based on the above embodiment, in one embodiment, sample data of different time periods are obtained, dynamic indicator values are obtained, and the dynamic indicator values are used to represent the state of the object, specifically including: Collecting a first data sample of a first object at a first time, and calculating a first target type indicator value based on the first data sample; collecting a second data sample of the first object at a second time, and calculating a second target type indicator value based on the second data sample; Calculating a difference between the second target type index value and the first target type index value, and generating evaluation information of the first object based on the difference; The second time is after the first time.

[0094] Still taking the number of enterprises as an example, workload change assessment: First evaluation: City A’s x=5,000, with a score of 447 points.

[0095] Second evaluation: City A's x=6,000, recalculated score: y 新 =475 points.

[0096] Difference: 475−447=+28 points, indicating that the workload of City A has increased.

[0097] Based on the above embodiment, in one embodiment, the following processing method is adopted for various target type indicator values: If it is determined that each object includes multiple target types, the target type index value corresponding to each target type is calculated separately; the weighted sum of multiple target type index values of the same object is calculated to obtain the total index value of the object in the group.

[0098] Specifically, they include: Definition of target type: Clarify business needs: Based on the specific business scenario and analysis objectives, determine the target types to focus on. For example, when analyzing the development of a municipal district, target types may include economic development indicators (such as GDP and number of enterprises), social and livelihood indicators (such as educational resources and medical resources), and environmental indicators (such as air quality and green coverage).

[0099] Establish a target type system: Classify and organize target types into a hierarchical system. For example, economic development indicators can be further divided into sub-types such as industrial output value and service industry output value.

[0100] Data association and matching: Data integration: Collect data related to the target type from different data sources (such as government statistical departments, industry associations, corporate databases, etc.) and integrate them into a unified dataset.

[0101] Object-target association: Associate each object (e.g., a city district) with its corresponding data for multiple target types. Data can be matched and associated using the object's unique identifier (e.g., district code).

[0102] Determination of weights: Expert evaluation method: Invite experts in related fields to evaluate the importance of each target type based on their business experience and professional knowledge and determine the corresponding weights.

[0103] Analytic Hierarchy Process (AHP): By building a hierarchical model, the relative importance of different target types is compared and the weights are calculated.

[0104] Data-driven approach: Automatically learn the weight of each target type based on historical data and machine learning algorithms.

[0105] Apply the weighted sum index value to specific business scenarios, such as ranking, evaluation, and decision-making. For example, rank city districts based on their sum index values and identify well-developed districts for promotion.

[0106] In some embodiments, the multimodal survey data comprehensive analysis system may include multiple functional modules composed of computer program segments. The computer program of each program segment in the multimodal survey data comprehensive analysis system may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Function for comprehensive analysis of multimodal survey data.

[0107] In this embodiment, the multimodal survey data comprehensive analysis system can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0108] An acquisition module, configured to acquire a plurality of data samples, wherein the data samples include data of multiple modalities; An identification module is used to identify target type parameters from data samples using a multimodal recognition model and to count parameter values of the target type parameters in the data samples; The analysis module is used to generate a target type index value for each data sample in the overall data based on the parameter values of the multiple data samples and predefined calculation rules.

[0109] Figure 3 The multimodal survey data comprehensive analysis method provided for the embodiment of the present application can be applied to a device. Those skilled in the art will understand that the device structure involved in the embodiment of the present invention does not constitute a limitation on the device, and the device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In the embodiment of the present invention, the device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0110] The device 300 may include a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0111] The memory 320 can be used to store execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 can perform some or all of the steps in the above-described method embodiments.

[0112] The processor 310 is the control center of the storage device, which uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and / or processes data by running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 310 can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0113] The communication unit 330 is configured to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.

[0114] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0115] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code, and includes instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0116] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0117] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, and can be electrical, mechanical or other forms.

[0118] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0119] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0120] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, persons of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any changes or substitutions that can be easily conceived by persons skilled in the art within the technical scope disclosed in the present invention shall be within the scope of protection of the present invention.

Claims

1. A comprehensive analysis method for multimodal survey data, characterized in that: include: Acquire multiple data samples, where the data samples include data of multiple modalities; Using a multimodal recognition model to identify target type parameters from data samples, and counting parameter values of the target type parameters in the statistical data samples; According to the parameter values of the plurality of data samples and predefined calculation rules, a target type index value of each data sample in the overall data is generated.

2. The method according to claim 1, characterized in that The method further comprises: Collecting a first data sample of a first object at a first time, and calculating a first target type indicator value based on the first data sample; collecting a second data sample of the first object at a second time, and calculating a second target type indicator value based on the second data sample; Calculating a difference between the second target type index value and the first target type index value, and generating evaluation information of the first object based on the difference; The second time is after the first time.

3. The method according to claim 1, characterized in that Acquire multiple data samples, where the data samples include data of multiple modalities, including: A plurality of data samples corresponding to a plurality of objects are obtained, where the data samples include video data, text data, and image data.

4. The method according to claim 3, characterized in that The method further comprises: The data samples of multiple objects corresponding to the same parent object are divided into the same group, and the data samples in the same group are analyzed.

5. The method according to claim 1, characterized in that The multimodal recognition model is used to identify target type parameters from data samples, and the parameter values of the target type parameters in the statistical data samples are counted, including: Store the target type parameters identified by the multimodal recognition model from the data sample in a specified format; The stored target type parameters are deduplicated, and the parameter values of the deduplicated target type parameters are counted.

6. The method according to claim 1, characterized in that Generating a target type index value for each data sample based on the parameter values of the plurality of data samples and a predefined calculation rule, including: Predefined calculation rules: Where y is the target type indicator value of the object to which the data sample belongs; x is the parameter value of the target type parameter of the data sample; N is the variation factor, including the integer part and the decimal part. The integer part is the number of digits in the sum of the parameter values of the objects in the same group, and the decimal part is the first digit of the sum of the parameter values of the objects in the same group. The parameter values of the target type parameters of the multiple objects in the same group are substituted into the predefined calculation rule to obtain the target type index value of each object.

7. The method according to claim 6, characterized in that The method further comprises: If it is determined that each object includes multiple target types, the target type index value corresponding to each target type is calculated respectively; Calculate the weighted sum of multiple target type index values of the same object to obtain the total index value of the object within the group.

8. A multimodal survey data comprehensive analysis system, characterized by: include: An acquisition module, configured to acquire a plurality of data samples, wherein the data samples include data of multiple modalities; An identification module is used to identify target type parameters from data samples using a multimodal recognition model and to count parameter values of the target type parameters in the data samples; The analysis module is used to generate a target type index value for each data sample in the overall data based on the parameter values of the multiple data samples and predefined calculation rules.

9. A device, characterized in that include: Memory, used to store a comprehensive analysis program for multimodal survey data; A processor is configured to implement the steps of the multimodal survey data comprehensive analysis method according to any one of claims 1 to 7 when executing the multimodal survey data comprehensive analysis program.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a multimodal survey data comprehensive analysis program, which, when executed by a processor, implements the steps of the multimodal survey data comprehensive analysis method according to any one of claims 1 to 7.