Multimodal operator frameworks and methods, systems and storage medium for data processing
The multimodal operator framework addresses the challenge of processing and fusing multimodal data by simplifying data management and analysis, thereby enhancing efficiency and reducing operational complexity.
Patent Information
- Application Number
- PCT/CN2024/119112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-09-14
- Publication Date
- 2025-06-12
AI Technical Summary
Current big data analytics platforms lack the capability to systematically support associative analysis of multimodal data, leading to high costs and increased complexity in processing and fusing different modalities of data.
A multimodal operator framework is introduced, comprising source, processing, and output operators, which enables the processing and fusion of initial multimodal data within a multimodal data table, simplifying data management and analysis.
The framework reduces the complexity and operational threshold of multimodal data fusion and associative analysis, improving the efficiency of data processing and analysis while allowing data analysts to focus on core business logic.
Smart Images

Figure CN2024119112_12062025_PF_FP_ABST
Abstract
Description
MULTIMODAL OPERATOR FRAMEWORKS AND METHODS, SYSTEMS AND STORAGE MEDIUM FOR DATA PROCESSING
[0001] CROSS REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese Patent Application NO. 202311673814.7, filed on December 07, 2023, the entire contents of which are hereby incorporated by reference.TECHNICAL FIELD
[0003] The present disclosure relates to the field of data processing, and in particular, to a multimodal operator framework and a method, a system, and storage medium for data processing.BACKGROUND
[0004] With the development of 5G and IoT as well as the popularity of video surveillance devices, a large amount of data is constantly being produced, and the large amount of data includes audio, video, images, text, and other modalities. With multidimensional and multimodal data, it may be possible to realistically describe the full picture of an object from all angles, thus implementing a digital twin. To truly describe an object and realize the digital twin, it is necessary to realize the multimodal big data association and fusion analysis.
[0005] Traditional platforms for big data analysis tend to focus on analyzing structured data from business databases, as well as textual data such as logs, or on analyzing data from a particular modality or business. The associative analysis between multimodal data is not supported by a big data platform having a generalized and systematic capability and a large amount of cost may be needed by the associative process between a plurality of multimodal data. And there are big differences in the way different multimodal data are processed, which requires data analysts to master various multimodal data processing methods and further raises the threshold of fusion analysis of multimodal data.
[0006] Therefore, it is desired to provide a multimodal operator framework, and a method, a system, and storage medium for data processing, realizing a fusion analysis between a variety of different multimodal data, simplifying the difficulty of managing, developing, and analyzing multimodal data, reducing the complexity and operational threshold of the fusion and associative analysis between multimodal data, and improving the efficiency of multimodal data processing and analysis.SUMMARY
[0007] One or more embodiments of the present disclosure provide a multimodal operator framework wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator; the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data to the at least one processing operator; the at least one processing operator is configured to processing result by processing the initial multimodal data and send the processing result to the output operator; the at least one output operator is configured to add output result generated based on the processing result to the multimodal data table.
[0008] One or more embodiments of the present disclosure provide a method of processing data, wherein the method comprises: obtaining a multimodal data table including initial multimodal data; and inputting the multimodal data table into a multimodal operator framework for data processing to obtain processing result.
[0009] One or more embodiments of the present disclosure provide a data processing system, wherein the system comprises at least one processor and at least one memory; the at least one memory for storing computer instructions; the at least one processor for executing at least a portion of the computer instructions to implement: obtaining a multimodal data table including initial multimodal data; and inputting the multimodal data table into a multimodal operator framework for data processing to obtain processing result, wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator; the at least one source operator is configured to obtain initial multimodal data in the multimodal data table and input the initial multimodal data to the at least one processing operator; the at least one processing operator is configured to processing result by processing the initial multimodal data and send the processing result to the output operator; and the at least one output operator is configured to add output result based on the processing result to result the multimodal data table.
[0010] One or more embodiments of the present disclosure provide a data processing system, wherein the system comprises: an acquisition module for obtaining a multimodal data table, the multimodal data table comprising initial multimodal data; a processing module for inputting the multimodal data table into a multimodal operator framework for data processing to obtain processing result, wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator; the at least one source operator is configured to obtain initial multimodal data in the multimodal data table and input the initial multimodal data into the at least one processing operator; the at least one processing operator is configured to obtain processing result by processing the initial multimodal data and send the processing result to the output operator; and the at least one output operator is configured to add output result generated based on the processing result to the multimodal data table.
[0011] One or more embodiments of the present disclosure provide a computer-readable storage medium, the storage medium storing computer instructions, and when the computer reads the computer instructions in the storage medium, the computer performs a method of data processing comprising: obtaining a multimodal data table, the multimodal data table comprising initial multimodal data; inputting the multimodal data table into a multimodal operator framework for data processing to obtain processing result, wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator; the at least one source operator is configured to obtain initial multimodal data in the multimodal data table and input the initial multimodal data into the at least one processing operator; the at least one processing operator is configured to obtain processing result by processing the initial multimodal data and send the processing result to the output operator; and the at least one output operator is configured to add output result generated based on the processing result to the multimodal data table.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The present disclosure will be further illustrated by way of exemplary embodiments, which will be described in detail by means of the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbering denotes the same structure, wherein:
[0013] FIG. 1 is a schematic diagram illustrating an exemplary operator framework 100 according to some embodiments of the present disclosure;
[0014] FIG. 2 is a schematic diagram illustrating an exemplary storage structure for multimodal data in a multimodal data table according to some embodiments of the present disclosure;
[0015] FIG. 3 is a schematic diagram illustrating an exemplary operator framework 300 according to another embodiments of the present disclosure;
[0016] FIG. 4A is a schematic diagram illustrating an exemplary operator framework 410 according to some embodiment of the present disclosure;
[0017] FIG. 4B is a schematic diagram illustrating an exemplary operator framework 430 according to another embodiment of the present disclosure;
[0018] FIG. 4C is a schematic diagram illustrating an exemplary operator framework 450 according to another embodiment of the present disclosure;
[0019] FIG. 5 is a schematic diagram illustrating an exemplary data processing device according to some embodiments of the present disclosure;
[0020] FIG. 6 is a diagram illustrating exemplary modules of a data processing system shown in accordance with some embodiments of the present disclosure;
[0021] FIG. 7 is a flow chart illustrating an exemplary process of processing data according to some embodiments of the present disclosure;
[0022] FIG. 8 is a schematic diagram illustrating an exemplary computer-readable storage medium according to some embodiments of the present disclosure.DETAILED DESCRIPTION
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the description of the embodiments are briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and it is possible for a person of ordinary skill in the art to apply the present disclosure to other similar scenarios in accordance with these drawings without creative labor. Unless obviously obtained from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.
[0024] It should be understood that the terms "system" , "device" , "unit" and / or "module" as used herein are used as a way to distinguish between different components, elements, parts, sections, or assemblies at different levels. However, the words may be replaced by other expressions if other words accomplish the same purpose.
[0025] As shown in the present disclosure and in the claims, unless the context clearly suggests an exception, the words "a" , "one" , "a" and / or "the" do not refer specifically to the singular and may include the plural. Generally, the terms "including" and "comprising" suggest only the inclusion of clearly identified steps and elements, and these steps and elements do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0026] Flowcharts are used in the present disclosure to illustrate operations performed by a system according to embodiments of the present disclosure. It should be appreciated that the preceding or following operations are not necessarily performed in an exact sequence. Instead, steps may be processed in reverse order or simultaneously. Also, it is possible to add other operations to these processes, or to remove a step or steps from these processes.
[0027] The purpose of data analysis is to centralize and distill the information hidden in a large amount of seemingly haphazard data, so as to find out the intrinsic laws of the object under study. In practice, data analysis helps people make judgments in order to take appropriate actions. Data analysis is the process of collecting data in an organized and purposeful manner, analyzing it and making it into information.
[0028] And with the increasing development of artificial intelligence technology, multimodal data fusion and analysis techniques have been widely used in the field of artificial intelligence. The multimodal data refers to data from multiple data sources which may be visual, auditory, texture, etc., and which are often present in a scene simultaneously. The fusion and analysis of the multimodal data are the fusion of disparate data from different data sources to obtain more accurate and comprehensive information to help people better understand and use the data.
[0029] Current data analytics methods or devices typically focus on analyzing structured data from business databases or text-based data such as logs, or on analyzing data for a particular modality or business. Big data analytics platforms lack the ability to generalize and systematically support the associative analysis of multimodal data, making it more costly for data analytics and data developers to deal with associative processing between different modalities and unable to focus their energy on the core business. And, due to the large differences in the processing of different modal data, data analysts often need to master a variety of modal data processing methods before they may perform fusion analysis of the multimodal data, resulting in a huge waste of manpower and further increasing the cost of fusion analysis of the multimodal data.
[0030] According to some embodiments of the present disclosure, a method, a system and storage medium for multimodal data processing may be determined based on a big data processing framework, a data middle-ground, and a data lake, realizing a systematic framework for the common management of structured data and a plurality of unstructured data of multimodality, and proposing a set of operator frameworks for multimodal data analysis based on the systematic framework, to solve the convenience problems of multimodal data management, multimodal data development and analysis (especially the fusion and associative analysis between multiple different modal data) . Through the built-in operators and algorithms for a plurality of modalities and scenarios and the graphical drag-and-drop development interface, it simplifies the difficulty of analyzing unique modal data, reduces the complexity of the multimodal data fusion analysis and the threshold of operation, and provides a convenient and efficient multimodal data analysis environment for the data analysts and developers, allowing to focus on the core business logic of data analysis and significantly improving the efficiency of multimodal data analysis.
[0031] FIG. 1 is a schematic diagram illustrating an exemplary operator framework 100 according to some embodiments of the present disclosure.
[0032] As shown in FIG. 1, the operator framework 100 includes a source operator 110, a processing operator 120, and an output operator 130. The source operator 110 is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data into the processing operator 120; the processing operator 120 is configured to process the initial multimodal data to obtain processing result and send the processing result to the output operator 130; the output operator 130 is configured to add the processing result to the multimodal data table.
[0033] An operator refers to a computation unit in the operator framework 100 or a code module of code used to implement specific operator logic. For example, an operator may be used to perform text analysis, translation, sentiment analysis, or the like. Operators in the operator framework 100 may be understood as nodes in a processing flow and each operator may represent a node performing a step in the processing.
[0034] The operator framework 100 is a collection of a plurality of operators for managing, fusing, analyzing, or the like of data. For example, the operator framework 100 may be used to perform reading and writing, storing, processing, analyzing, or the like, or a combination thereof, of data. In some embodiments, individual operators in the operator framework 100 may be used to process structured data and unstructured data. More descriptions of the structured data and the unstructured data may be found elsewhere in the present disclosure.
[0035] The source operator 110 serves as the starting point of the operator framework 100. The source operator 110 is configured to obtain initial multimodal data in the multimodal data table and output the initial multimodal data to the processing operator 120. It should be noted that FIG. 1 merely shows single one source operator, but the count of source operators of the operator framework 100 may exceed 1. Each the source operator may correspond toa multimodal data table and configured to obtain initial multimodal data from the corresponding multimodal data table.
[0036] The multimodal data table refers to a data table used to store multimodal data.
[0037] The multimodal data refers to the data in multiple different modalities. The multimodal data may include data from a plurality of data sources, for example, a multimodal data table can include multiple modal data such as picture data, sound data, video data, text data, and so on.
[0038] In some embodiments, the multimodal data includes the structured data, the unstructured data of a single modality, the unstructured data of multimodality, or a combination thereof.
[0039] The structured data refers to the data that may be represented and stored using a relational database and expressed in a two-dimensional form. Characteristics of the structured data include that the structured data is in rows, a row of data represents information about an entity, and the attributes of each row of data are the same. For example, the structured data may include data in a MySQL database, a csv file, etc. As another example, the structured data may include data of the types including integer, floating-point, fixed-point, date or time, string, text, etc.
[0040] The above relational database may include an enterprise ERP system, a financial system, a medical HIS database, or the like.
[0041] The unstructured data refers to the data that is unfit for the two-dimensional logical table in a relational database. The unstructured data may not be organized in a pre-defined data model or a pre-defined way. For example, the unstructured data may include an office document, a Pdf document, an email, a log, an XML format file, a JSON format file, an image, an audio, a video information, etc.
[0042] In some embodiments, the unstructured data includes the unstructured data of a single modality and / or the unstructured data of multimodality.
[0043] The modality of the unstructured data refers to the data format of the unstructured data, and each data format corresponds to a modality. For example, data formats include audio, video, text, etc.
[0044] The unstructured data in a single modality refers to the unstructured data that consists of data in a single data format.
[0045] The unstructured data of multimodality refers to the unstructured data that contains data in multiple data formats.
[0046] The initial multimodal data refers to unprocessed multimodal data in the multimodal data table.
[0047] In some embodiments, the initial multimodal data may be all the data contained in the multimodal data table. In other embodiments, the initial multimodal data may be data of a portion contained in a multimodal data table.
[0048] Processing of the multimodal data includes, but is not limited to, operations such as data conversion, data extraction, data filtering, etc. More descriptions for data processing of the multimodal data may be found elsewhere in the present disclosure.
[0049] In some embodiments, the multimodal data table may include both structured data and unstructured data, but the multimodal data table is still a structured data table.
[0050] In some embodiments, the multimodal data table may include descriptive information for each of the multimodal data. The descriptive information of a piece of data may include a field name, a field type, and a key content description, or the like, or a combination thereof of the piece of data.
[0051] The field name is a field used to identify the data. Field names corresponding to the same type of data may be the same. Field names are specified by the user and follow certain naming rules in different systems. For example, for data that represents a “name” , the field name may be “name” or “xm” (i.e., the initials of the Chinese pinyin “xing ming” for "name" ) . For data used to represent the “ID number” , the field name may be “ID number” or sfzhm (i.e., the initials of the Chinese pinyin “shen fen zheng hao ma” for "ID number" ) .
[0052] The field type refer to a field used to describe the modal type of data. The field type include a structured field type or an unstructured field type.
[0053] The structured field type correspond to the structured data, which may be defined by the system. In some embodiments, the structured field type may include a type such as integer, floating point, fixed point, date, time, string, text, etc. For example, the structured field type may include an int type, a string type, or the like. The int type indicates that the structured data is an integer and the string type indicates that the structured data is a string.
[0054] The unstructured field type correspond to the unstructured data. The unstructured field type may be specified by the user. In some embodiments, the unstructured field type may include a type such as image, video, audio, or the like. For example, the field type of the picture data may be of Picture type, the field type of the video data may be of Video type, etc., and the field type of the audio data may be of Audio type.
[0055] It should be noted that the field types in the multimodal data table are only used to characterize the essence of the data type, and do not qualify the actual file format of the unstructured data such as Picture, Video, Audio, etc. For example, the actual file format of the Picture type of data may be any type of picture format such as jpg, png, etc.; the actual file format of the Video type of data may be any type of video format, such as mp4, avi, etc.; the actual file format of the Audio type of data may be any type of audio format, such as mp3, wav, etc.
[0056] The key content description is an associated description of the key content in the data. For example, the key content description may include the data type, the range of values, the practical significance, etc., of the data. For example, the key content description of the data indicating the gender may be 0 or 1, where 0 indicates that the gender is male, and 1 indicates that the gender is female, etc.
[0057] For example, an exemplary multimodal data table bout people information may be shown in Table. 1.
[0058] Table. 1:
[0059] In this case, there are three fields in Table. 1, which are field name, field type, and key content description. For the row of data with the field name id_num, the field type is string, so the row is structured data. The key content description indicates that the row of data with the field name id_num is the person's ID number. For the row of data whose field name is picture, the field type is of Picture type, so the row of data is unstructured modal data, and the key content description indicates that the data is the person's recent Picture.
[0060] In some embodiments, the real data of the multimodal data may be stored in the multimodal data table.
[0061] In some embodiments, the real data of the multimodal data (e.g., raw content of the type of images, video, audio, etc. ) may be stored in a memory external to the multimodal data table for the storage efficiency, storage performance, etc.
[0062] When the real data of the multimodal data is stored in the external memory, the actual storage address of the real data of the multimodal data may be stored in the multimodal data table, and the user may access the real data of the multimodal data through the actual storage field to realize reading and writing the multimodal data.
[0063] In some embodiments, the multimodal data table is configured to store the metadata for the multimodal data. The metadata includes at least one of a modal type field or an actual storage field. The modal type field is used to describe a file type of the multimodal data, and the actual storage field is used to describe an actual storage address of the real data of the multimodal data.
[0064] The metadata for the multimodal data refers to the data related to the multimodal data. For example, the metadata comprises data for describing a data type, a location, etc., of the multimodal data. The metadata is different from the real data.
[0065] In some embodiments, the metadata of the multimodal data may be binary data.
[0066] In some embodiments, the metadata of the multimodal data may include at least one of a modal type field or an actual storage field.
[0067] The modal type field is used to describe the file type of the multimodal data. In some embodiments, the file type includes structured data or unstructured data. The file type may be further subdivided. For example, the file type of the unstructured data may be subdivided to include an image type, a video type, an audio type, etc.
[0068] The actual storage field is used to describe the actual storage address of the real data of the multimodal data. For example, the actual storage field may describe the file ID or HTTP URL of the actual storage address, etc. That is, the actual storage field may indicate where the real data of the multimodal data is stored, so that it may be quickly located when looking up the data.
[0069] In some embodiments, the actual storage addresses of the multimodal data for different modalities may be the same or different.
[0070] In some embodiments, the multimodal data of different modalities may correspond to different storage engines, whose storage addresses may be accessed via HTTP or a proprietary SDK.
[0071] In some embodiments, when the real data of the multimodal data is stored in an external memory (e.g., an external distributed file system) outside of the multimodal data table, the multimodal data table may only include the metadata of the multimodal data.
[0072] In some embodiments, when the real data of the multimodal data is stored in an external memory, reading and writing of the real data of the multimodal data may be realized by a read connector and a write connector. More descriptions of the read connector and the write connector may be found elsewhere in the present disclosure.
[0073] In other embodiments, when the real data of the multimodal data is stored in an external memory outside of the multimodal data table, the metadata of the multimodal data may also be stored in an external memory outside of the multimodal data table. In this case, the multimodal data table may include only the identifiers of the metadata. An identifier of the metadata may include an actual storage address of the metadata to facilitate checking the metadata.
[0074] In some embodiments, the metadata of the multimodal data includes the actual storage manner, a link parameter, and / or other descriptive information (e.g., data format, etc. ) .
[0075] The actual storage manner is the way the real data of the multimodal data is stored. For example, the real data of the multimodal data may be stored in a type such as binary, bytes, or the like.
[0076] A link parameter may be used to access the external memory of the real data of the multimodal data, e.g., an user name for access, an user password, etc.
[0077] According to some embodiments of the present disclosure, the multimodal data table is configured to store the metadata of the multimodal data, and the real data of the multimodal data is stored in an external memory, which may effectively improve the storage efficiency of the data and the storage performance of the multimodal data table.
[0078] The multimodal data table may be constructed in a number of ways.
[0079] In some embodiments, the processor may obtain the multimodal data table based on a native data management system The native data management system includes a database management system (e.g., Database Management System) , a data warehouse (e.g., a data warehouse implemented based on Hive) , or other database / data warehouse-like systems that support the management and use of the structured data in terms of table concepts, e.g., Presto, Trino, etc.
[0080] The native data management system is typically designed for the structured data and not very good at storing and managing the unstructured data. While there are field types such as binary, bytes, blob, etc. defined that allow for the storage of binary data, there is no direct perception of its data modality, and only very limited general-purpose operations are available, e.g., no direct support for image recognition, conversion, etc. In addition, the native data management system is not as efficient as a distributed file system for storing the unstructured data such as video, images, audio, etc.
[0081] According to some embodiments of the present disclosure, the processor may obtain the multimodal data table based on the native data management system. For example, the processor may extend the original data in the native data management system to obtain the multimodal data table. In some embodiments, the processor may add descriptive information for each type of modal data, which in turn may be constructed to obtain the multimodal data table.
[0082] FIG. 2 is a schematic diagram illustrating an exemplary storage structure for multimodal data in a multimodal data table according to some embodiments of the present disclosure.
[0083] For example, the processor may extend the original MySQL database table to get a multimodal data table. As shown in FIG. 2, a raw MySQL database table includes two integer types of data and two string types of data. The processor adds image data, video data, and audio data by expanding the original data types to get a table that includes two integer types, two string types, one image type, one video type, and one audio type in a multimodal data table. The structured data includes the data of integer type and string type, stored in a MySQL database. The unstructured data may include the data of photo type and the data of audio type stored in a distributed system (e.g., a hdfs system) , and the data of video type stored in other storage systems (e.g., a Ceph system) .
[0084] Taking the photo field as an example, the metadata included in the photo field may include the modal type: "Picture" , the storage manner "hdfs, " the storage path: "hdfs: / / nn1. example. com / file1" , the account for accessing the database: "test" and the password for accessing the database "test@123456" , or the like.
[0085] In some embodiments, the multimodal data table may be constructed by designing data management system or a framework. For example, the processor may define the corresponding descriptive information or the metadata, etc., for the structured data and the unstructured data, respectively, which may be constructed to obtain the multimodal data table.
[0086] Through the setup of the multimodal data table, the unstructured modal data may be better managed, and the data modality of the stored unstructured data may be directly perceived, which facilitates the subsequent data analysis or data fusion.
[0087] The processing operator 120 may be a downstream operator of the source operator 110. The processing operator 120 may be an operator in the operator framework 100 for processing data.
[0088] In some embodiments, the processing operator 120 may be configured to process the initial multimodal data from the source operator 110 to obtain processing result and send the processing result to the output operator 130.
[0089] In some embodiments, the processing operator 120 may be configured to process the non-initial multimodal data to obtain processing result and send the processing result to the output operator 130. In some embodiments, the non-initial multimodal data may be the processed multimodal data. For example, the non-initial multimodal data may be a processing result from an upstream processing operator. Detailed descriptions of the multilevel processing operators may be referred to in relevant descriptions of FIG. 3. In some embodiments, the non-initial multimodal data may be the processing result from an upstream sorting operator or a transaction management operator, see below for more descriptions of the sorting operator, and the transaction management operator.
[0090] In some embodiments of, the processing operator 120 may perform data processing on the unstructured data as well as the structured data. For example, the processing operator 120 may perform data processing on the structured data that may not be processed by the native data management system. In some embodiments, for the structured data of common types may be processed by the native data management system (e.g., the structured data of common types may include the int, the string, the timestamp, etc. ) , the processing operator 120 may do no processing thereof and directly passthrough to the output operator 130.
[0091] The data processing performed by the processing operator 120 includes operations such as data conversion, data computation, data filtering, etc.
[0092] In some embodiments, the data conversion includes conversion of data in the same modality (e.g., from pictures to pictures) , conversion from the unstructured data to the structured data (e.g., the conversion may be to convert the unstructured data such as pictures to the structured data such as numbers, strings, etc. ) , conversion from the unstructured data in one modality to the unstructured data in other modalities (e.g., the conversion may be to extract audio from video) , etc., or any combination thereof. In some embodiments, the processing operator 120 may perform multiple data transformations, for example, extracting various data such as the structured data, video clips, images, etc., from a video simultaneously.
[0093] Data computation refers to an arithmetic operation, a relational operation, a logical operation, or the like, or a combination thereof on the multimodal data.
[0094] Data filtering is the operation of filtering, excluding, or extracting a specific part of a data set. Data filtering may be performed based on a specific condition or a rule to select eligible data from a dataset to meet a specific need or for further analysis. In some embodiments, the data filtering includes row-level filtering, column-level filtering, conditional filtering, text filtering, temporal filtering, unique value filtering, advanced filtering, etc., or a combination thereof.
[0095] In some embodiments e, the processor may implement data filtering via a filter instruction. For example, a filter instruction may be expressed as “select (field name) from (multimodal data table name) where (query condition) . ” For example, the multimodal data table "name_age" consists of the names (the corresponding field name is “name” ) and ages (the corresponding field name is “age” ) of the personnel, if you want to query the names of the personnel whose age is greater than 20 years old, the filtering command may be “select (name) from (name_age) where (age>20) . ”
[0096] In some embodiments, for the structured data, the query conditions may be determined directly based on the values taken from the structured data.
[0097] In some embodiments, for the unstructured data, the processor may determine a query condition based on feature values of the unstructured data. For example, the processing operator may determine a characteristic value of the unstructured data, such as an image, a video, an audio, or the like, and the feature value of the unstructured data t obtained by filtering may be used as a query condition.
[0098] In some embodiments, for the filtering, computing, converting, and other operations of the unstructured data, the processor may construct or insert a corresponding plug-in in the processing operator 120, and the f plug-in may be implanted via a dynamic link library for implementation of language, such jar packages, c / c++, and other language, python scripts, etc.
[0099] for the multimodal data table that is determined based on the native data management system, building the above plug-in should first follow the extended interface and constraint of the native data management system, such as defining plug-in as a custom function of the native data management system. In some embodiments, the operation of the plug-in may either rely on the computational power of a big data computing framework such as Flink, Spark, or a database such as Doris, or it may be based on a completely independent, self-developed service or frameworks to run.
[0100] In some embodiments, the processing result of the processing operator 120 may include a result obtained after performing at least one of data conversion, data computation, data filtering, or the like on the initial multimodal data.
[0101] In some embodiments, the processing result of the processing operator 120 includes data of another modalities converted from the data of each modality in the initial multimodal data, data of one or more specified modalities extracted from the initial multimodal data, data of one or more specified modalities filtered from the initial multimodal data, or the like, or a combination thereof.
[0102] In some embodiments, with respect to the initial multimodal data is the unstructured data of multimodality, the multimodal data of another modalities converted from the initial multimodal data may include: the structured data converted from the unstructured data, the unstructured data of another modality and the structured data converted from the unstructured data of one modality, the unstructured data and the structured data of another modality converted from the unstructured data of one modality, or the like, or a combination thereof. For example, the structured data converted from the unstructured data may include the structured data of type int or string converted from the unstructured data of type picture. For example, the structured data converted from the unstructured data of one modality and the unstructured data of the another modality converted from the unstructured data of the one modality may include the unstructured data of another picture type converted from the unstructured data of the picture type.
[0103] In some embodiments, with respect to the initial multimodal data is the structured data, the multimodal data of another modalities converted from the initial multimodal data may include structured data of another modality converted from the structured data of one modality.
[0104] In some embodiments, the data of the specified modality extracted from the initial multimodal data may include unstructured data of a single modality extracted from the initial multimodal data, unstructured data of multimodality extracted from the initial multimodal data, structured data extracted from the initial multimodal data, etc., or a combination thereof. For example, the multimodal data extracted from the initial multimodal data of the remaining modalities may include an extraction of frames in the video to obtain pictures, a culling of audio in the video data to obtain silent video, and also extracting structured data, video clips, pictures, etc., from the video simultaneously.
[0105] In some embodiments, the data of one or more specified modalities filtered from the initial multimodal data may include unstructured data of a single modality filtered from the initial multimodal data, unstructured data of multimodality filtered from the initial multimodal data, structured data filtered from the initial multimodal data, etc., or a combination thereof.
[0106] It should be noted that FIG. 1 merely shows only one processing operator and in some embodiments, more than one processing operators may be provided in the operator framework 100.
[0107] In some embodiments more than one processing operators 120 in the same operator framework 100 may be paralleled as downstream operators of the source operator 110. In this embodiment, each processing operator of the more than one processing operators corresponds to initial multimodal data.
[0108] In some embodiments, the more than one processing operator in the same operator framework 100 may be connected in series as downstream operators of the source operator 110 and form multilevel processing operators. Detailed descriptions of the multilevel processing operators may be referred to in relevant descriptions of FIG. 3.
[0109] In some embodiments, the operator framework 100 may include a conversion operator, a computation operator, a filtering operator, an extraction operator, etc., or a combination thereof. The conversion operator is used to perform a processing operation for data conversion. The computation operator is used to perform a processing operation of data computation. The filtering operator is used to perform a processing operation for data filtering. The extraction operator is used to perform the processing operation of data extraction.
[0110] In some embodiments, when one or more of the conversion operator, the computation operator, the filtering operator, the extraction operator, etc., are provided in the operator framework 100, these operators may be connected in parallel or in series as the downstream operators of the source operator 110.
[0111] The output operator 130 serves as the endpoint of the operator framework 100. The output operator 130 may be a downstream operator of the processing operator 120. The output operator 130 may be an operator in the operator framework 100 for performing data output.
[0112] In some embodiments, the output operator 130 is configured to send a processing result to a multimodal data table. It should be noted that FIG. 1 merely shows only one output operator but in some embodiments, more than one output operator may be provided in the operator framework 100, each output operator 130 being used to write processed multimodal data to a multimodal data table. In some embodiments, the multimodal data table corresponding to the output operator 130 may be the multimodal data table corresponding to the initial multimodal data or may be a new multimodal data table.
[0113] In some embodiments, at least one source operator 110, at least one processing operator 120, and at least one output operator 130 may have been provided in the operator framework 100. Embodiments of the present disclosure do not limit the count of the various types of operators in the operator framework 100. For example, the operator framework 100 may include any number of the source operators 110, the processing operators 120, and the output operators 130 according to user requirements or business requirements. In the data processing device, the user may arrange the operator framework 100 by graphically displaying the operator framework page and dragging the various operators and also arrange the operator framework in the form of a text-input SQL statement or a SQL-like statement, without qualification.
[0114] Depending on the business needs, each of the operators in the above operator framework may be combined arbitrarily and repetitively; the multimodal result table may be queried and utilized in the upper tier applications including but not limited to displaying the image content and playing audio / video.
[0115] In some embodiments, the processor may preset or define different source operators 110, processing operators 120, and output operators 130 for processing the initial multimodal data of different modalities so that the initial multimodal data may be verified at the time of inputting it into the source operators 110, the processing operator 120, and the output operator 130, the processor may verify the modalities of the initial multimodal data. The input initial multimodal data is processed only if the input initial multimodal data is of a modality supported by the source operator 110, the processing operator 120, and the output operator 130. For example, the user may define a processing operator 120 for processing the initial multimodal data of a picture type, and if the input to the processing operator 120 is verified to be the initial multimodal data of a video type, the processing operator 120 refuses to perform processing the initial multimodal data of a video type.
[0116] In some embodiments, the user may graphically display the process of constructing the operator framework and drag and drop the specific operator to form an unified multimodal data analysis and processing framework for realizing the filtering, processing, association, and result output of the multimodal data.
[0117] In some embodiments, the user may customize the operator that extends its various other functions or the like as desired.
[0118] According to some embodiments of the present disclosure, a multimodal operator framework is provided that unifies the unstructured data and the structured data into a multimodal data table, realizes the common management of the structured data and a plurality of unstructured data of multimodality, and solves the convenience problem of multimodal data management, multimodal data development and analysis (especially the fusion and associative analysis between a variety of different modal data) . At the same time, through the built-in multimodal, multi-scenario operation operators and algorithms, as well as graphical drag-and-drop development interface, to simplify the difficulty of analyzing unique modal data, reduce the complexity of the fusion of the multimodal data analysis and the threshold of the operation, it provides a convenient and efficient multimodal data analysis environment for data analysts and data developers to focus on the core business logic of data analysis and significantly improve the efficiency of the multimodal data analysis. For example, users may read, write, query, and / or process unstructured data of multimodality directly in the usual way as they do with ordinary structured data, and the operator framework has a wider range of applications and higher efficiency.
[0119] In some embodiments, in order to enable the source operator 110 and the output operator 130 in the operator framework 100 to read and write the real data of the multimodal data, the operator framework 100 further includes a read connector and / or a write connector. The read connector is configured to obtain the real data of the initial multimodal data from the actual storage address of the initial multimodal data and send the real data of the initial multimodal data to the source operator 110; and the write connector is configured to write the real data of the target multimodal data output from the output operator 130 to the actual storage address of the initial multimodal data. The target multimodal data is the output of the output operator 130.
[0120] When the real data of the multimodal data is stored in the external memory, the processor may read and send the real data of the multimodal data to the source operator 110 via the read connector, the read connector may be set upstream at the upstream position of the source operator 110.
[0121] When the real data of the multimodal data is stored in an external memory, it is necessary to write the real data of the target multimodal data of the output operator 130 to the external memory via the write connector, the write connector may be provided at the downstream position of the output operator 130. In some embodiments, the metadata of the multimodal data is included in the multimodal data table, when the real data of the multimodal data is read via the read connector or written via the write connector, the metadata of the multimodal data is obtained from the multimodal data table, and based on a link parameter in the metadata of the multimodal data, a link is made to an actual storage address where the real data of the multimodal data is stored, thereby realizing the reading or writing of the real data of the multimodal data.
[0122] The actual storage manner in the metadata of the multimodal data determines which read connector is used to access the real data of the multimodal data, or determines which write connector is used to write the real data of the multimodal data. In some embodiments, the read connector and the write connector comprise a plurality of types, and the different types of the read connector and the write connector are used to connect to external memories corresponding to different actual storage manners. For example, when the real data of the multimodal data is stored in an hdfs system, a read connector and a write connector need to be defined for hdfs type of stored data. For example, the way of defining the read connector and the write connector may include obtaining a corresponding read connector and a write connector based on the hdfs client package.
[0123] The target multimodal data refers to processed multimodal data. For example, the target multimodal data may be the multimodal data that is obtained after operations such as data conversion, data computation, data filtering, etc.
[0124] In some embodiments, the real data of the target multimodal data may be binary data.
[0125] In some embodiments, the read connector or the write connector may read or write the multimodal data based on feature values of the multimodal data.
[0126] The feature values of the multimodal data refer to data that describe the relevant characteristics of the multimodal data. Feature values may be expressed as a feature vector. In some embodiments, the feature values of the multimodal data may be determined by any feasible algorithm such as a feature extraction algorithm.
[0127] In some embodiments, the processor may pre-build a vector database for storing feature vectors of the multimodal data and actual storage fields corresponding to the feature values. When the read connector read the multimodal data, the reading connector may simultaneously determine a feature vector satisfying a first preset condition from the vector database, read the real data of the multimodal data from the external memory according to the actual storage field corresponding to the feature vector satisfying the first preset condition. When the write connector writes the multimodal data, the actual storage field of the real data of the multimodal data may be associated with the feature value of the multimodal data and written to the vector database. In some embodiments, the first preset condition may include the vector distance satisfying a distance threshold, the vector distance being minimal, or the like.
[0128] According to some embodiments of the present disclosure, the processor decouples the operator (data analysis and the actual processing engine) and the actual storage address (the storage engine) through a read connector and a write connector to support a plurality of different processing and storage engines.
[0129] In some embodiments, the operator framework 100 may further include an association operator. The association operator is configured to associate any two of output result from the source operator 110 and / or output result from the processing operator 120.
[0130] The association operator is used to associate any two data contents so that their outputs are associated on one or more fields.
[0131] In some embodiments, the output result of the association operator may be used as an input of the output operator 130 or as an input of the processing operator 120.
[0132] In some embodiments, the association operator may be configured to associate any two outputs from the source operator 110. As shown in FIG. 4A the operator framework 410 includes two source operators 110 at upstream and the processing operator 120 at downstream, and the association operator 420 may associate two output result from the source operators 110 to generate an association result and input the association result (i.e., the output result of the association operator) to the processing operator 120 at downstream.
[0133] In some embodiments, the association operator may be configured to associate any two processing results from the processing operators. As shown in FIG. 4B, the operator framework 430 includes two processing operators 120, upstream and the output operator 130downstream, and the association operator 440 may associate two processing results from the processing operators 120 to generate an association result and input the association result (i.e., the output result of the association operator) to the output operator 130downstream. Different upstream processing operators 120 may correspond to different source operators 110, as shown in FIG. 4B. In some embodiments, the upstream different processing operators 120 may correspond to the same source operator 110, and there is no restriction here.
[0134] In some embodiments, the association operator may be configured to associate an output result from the source operator 110 with a processing result from the processing operator 120. As shown in FIG. 4C, the operator framework 450 includes a source operator 110 and a processing operator 120, upstream and an output operator 130downstream, and the association operator may associate an output result from the source operator 110 with a processing result from the processing operator 120 to generate an association result and input the association result (i.e., the output result of the association operator) to the output operator 130downstream. As shown in FIG. 4B, one of the outputs from the source operator 110 may be input into the processing operator 120 to obtain a processing result and the processing result outputted by the processing operator 120 may be input into the association operator 460 along with the outputs from the other source operator 110 in which the association is performed.
[0135] In some embodiments, the user may set the position of the association operator in the overall operator framework according to the actual needs, to choose to use the output result of the association operator as the output result of the overall operator framework or as the input of the output operator or the processing operator. That is to say, the user may specify the associated fields and the form of the association using the association operator based on the actual requirements. For example, the user may graphically display the output result of the source operator and / or the fields of the processing result of the processing operator, and drag and drop the fields to customize the arrangement of the association operator to associate the associated fields and the form of association.
[0136] According to some embodiments of the present disclosure, the output result of the association operator may be used as the input of the output operator or the input of the processing operator, which may effectively increase the degree of freedom of the operator framework to perform data processing, and the user may adjust the setting of the position of the association operator to adequately process the data.
[0137] In some embodiments, the association operator may associate the different multimodal data based on the feature values of the multimodal data. For example, the association operator may determine, based on the feature values of the multimodal data (e.g., the initial multimodal data or the non-initial multimodal data) output by the two upstream operators, two feature vectors that meets a second preset condition from a vector database, read the real data of the two multimodal data from the external memory according to the actual storage field corresponding to the two feature vectors, and thus associate the real data of the multimodal data output by the two upstream operators. In some embodiments, the second preset condition may include the vector distance of the two feature vectors satisfying a distance threshold, the vector distance being minimized, or the like.
[0138] According to some embodiments of the present disclosure, the output result of two source operators, the output result of any one source operator and the processing result of any one processing operator, or the processing result of two processing operators may be used as inputs to the association operator, which may improve the flexibility of the operator framework to perform association operations on data, improve the efficiency of association operations, and reduce unnecessary data processing.
[0139] In some embodiments, the operator framework 100 may further include a sorting operator. The sorting operator is configured to sort the multimodal data.
[0140] In some embodiments, the upstream operator of the sorting operator may be the source operator 110, and the downstream operator of the sorting operator is the processing operator 120. In this embodiment, the sorting operator may be used to sort the initial multimodal data obtained by the source operator 110, and send the initial multimodal data outputted by the source operator 110 to the processing operator 120 based on the sorting result.
[0141] In some embodiments, the upstream operator of the sorting operator is the processing operator 120 and the downstream operator of the sorting operator is the output operator 130. In this embodiment, the sorting operator is used to sort the processing result output by the processing operator 120 and send the output result of the processing operator 120 to the output operator 130 based on the sorting result.
[0142] In some embodiments, the upstream operator of the sorting operator is an upstream processing operator and the downstream operator is a downstream processing operator. In this embodiment, the sorting operator is used to sort the processing result of the upstream processing operator, and send the processing result of the upstream processing operator to the downstream processing operator based on the sorting result.
[0143] In some embodiments, the upstream operator of the sort operator may be an association operator and the downstream operator of the sort operator may be the processing operator 120. In the embodiment, the sorting operator may be used to sort the association result of the association operator output, and send the association result of the association operator output to the processing operator 120 according to the sorting result.
[0144] In some embodiments, the upstream operator of the sort operator may be an association operator and the downstream operator of the sort operator is the output operator 130. In the embodiment, the sorting operator may be used to sort the association result of the association operator output, and send the association result of the association operator output to the output operator 130 according to the sorting result.
[0145] In some embodiments, the sorting operator may implement sorting based on the feature values of the multimodal data. Sorting of the feature values may be achieved by a sorting algorithm such as a cluster analysis algorithm, an algorithm by vectorizing the feature values to determine similarity, or the like.
[0146] Sorting operators may also implement sorting in other ways. For example, the processor may implement sorting of the multimodal data based on business needs or user query statements, etc.
[0147] According to some embodiments of the present disclosure, by sorting the multimodal data using a sorting operator, the data processing efficiency of the operator framework is improved, and the output multimodal data table is more organized and more readable.
[0148] In some embodiments, the operator framework 100 may further include a sorting operator and an association operator.
[0149] In some embodiments, a transaction management operator may also be included in the operator framework 100. The transaction management operator is configured to perform single program management of the data processing flow. Single-program management means that for a piece of data, only one processing program may be performed to process the piece of data at the same time. For example, if processing program A is performed to process data X, and at this time, processing program B also is used to process data X as well. The transaction management operator may prevent processing program B from running until processing program A has finished, and then allows processing program B to run.
[0150] The transaction management operator may be located anywhere in the operator framework 100. The transaction management operator may also be located in other positions, which are not limited herein.
[0151] In some embodiments, the downstream operator of the transaction management operator may be the input operator 110. The transaction management operator maybe used for a single program management of the process of the initial multimodal data input operator 110.
[0152] In some embodiments, the downstream operator of the transaction management operator may be a processing operator 120. The transaction management operator can be used for a single program management of the processing operator 120 processing initial multimodal data.
[0153] In some embodiments, the downstream operator of the transaction management operator may be the output operator 130. The transaction management operator is used for single program management of the process of writing data by the output operator 130.
[0154] In some embodiments, the downstream operator of the transaction management operator maybe an association operator. The transaction management operator is used for single program management of the process of data association by the association operator.
[0155] In some embodiments, the downstream operator of the transaction management operator may be a sort operator. The transaction management operator is used for single program management of the process of sorting the sort operator.
[0156] In some embodiments, the processor, by using the transaction management operator to manage the handlers, may avoid a single piece of data being processed by a plurality of handlers at the same time resulting in an error.
[0157] In some embodiments, the operator framework 100 may include at least two of a sorting operator, a transaction management operator, and an association operator.
[0158] FIG. 3 is a schematic diagram illustrating an exemplary operator framework 300 according to another embodiment of the present disclosure.
[0159] As shown in FIG. 3, the operator framework 300 includes a source operator 110, multilevel processing operators 310, and an output operator 130. For one of the multilevel processing operators 310: the processing operator may receive a first processing result from an upstream processing operator and process the first processing result to generate a second processing result, and / or process the initial multimodal data input by the source operator 110 to generate the second processing, and send the second processing result to a downstream processing operator.
[0160] The multilevel processing operator 310 is an operator structure that consists of a plurality of levels of processing operators.
[0161] As shown in FIG. 3, the multilevel processing operator 310 may include a first-level processing operator 310-1, a second-level processing operator 310-2, ......, an N-level processing operator 310-n, n is equal to the count of the processing operators. The upstream operator of the first-level processing operator 310-1 is the source operator 110, i.e., the input of the first-level processing operator 310-1 may be the initial modal data input by the source operator 110. The first-level processing operator 310-1 is the upstream operator of the second-level processing operator 310-2, i.e., the processing result of the first-level processing operator 310-1 may be input to the second-level processing operator 310-2. Sequentially and analogously, the processing result of the second-level processing operator 310-2 may be input to downstream processing operators until the N-level processing operator 310-n, which is in the last tier, inputs its processing result to the downstream output operator 130.
[0162] In some embodiments, the count of processing operators in the multilevel processing operator 310 may be set according to the business requirements and is not limited herein.
[0163] More descriptions of the contents of the source operators, output operators, processing operators, and how processing operators process data may be referred to relevant descriptions of FIG. 1.
[0164] In some embodiments, by setting up a multilevel processing operator in the operator framework, a situation in which multimodal data may need to be processed multiple times according to business requirements is considered. The input of a downstream processing operator in the multilevel processing operators may be the processing result of an upstream processing operator, and the initial modal data is processed multiple times through the processing operators of the plurality of levels, until the final-level processing operator outputs the final processing result. The structure of the multilevel processing operator enables the initial modal data to be processed sufficiently, and increasing or decreasing the levels of the processing operator according to business requirements also improves the effectiveness of the multilevel operator in processing the initial modal data.
[0165] One or more embodiments of the present disclosure provide a data processing device. The data processing method described in the above embodiments may be realized based on a data processing device.
[0166] The data processing device is described below in connection with FIG. 5. is a schematic diagram illustrating an exemplary data processing device according to some embodiments of the present disclosure.
[0167] In some embodiments, as shown in FIG. 5, the data processing device 500 comprises at least one memory 510 and at least one processor 520, the at least one memory 510 being used for storing computer instructions, and the at least one processor 520 executing the computer instructions or part of the instructions to realize a data processing method including obtaining a multimodal data table, inputting the multimodal data table into an operator framework for data processing to obtain a processing result. For a more detailed description of the data processing method see FIG. 7 and its associated description. The operator framework for data processing may be found elsewhere in the present disclosure (e.g., FIG. 1 -FIG. 3 and the descriptions thereof) .
[0168] One or more embodiments of the present disclosure further provide a data processing method. The data processing method may be implemented based on a data processing system. Detailed descriptions of the contents of the data processing method may be referred to relevant descriptions of FIG. 7.
[0169] The data processing system is described below in connection with FIG. 6. is a diagram illustrating exemplary modules of a data processing system shown in accordance with some embodiments of the present disclosure. As shown in FIG. 6, the data processing system 600 includes an acquisition module 610 and a processing module 620.
[0170] In some embodiments, the acquisition module 610 may be configured to acquire a multimodal data table. For a more detailed description of the multimodal data table see FIG. 1 and its related description.
[0171] In some embodiments, the processing module 620 may be configured to input the multimodal data table into the operator framework for data processing to obtain a processing result. Wherein, the operator framework includes: at least one source operator, at least one processing operator, and at least one output operator. See FIG. 1 -FIG. 3 and their related descriptions for a more detailed description of the operator framework.
[0172] In some embodiments, the processing module 620 may be configured that the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data to the at least one processing operator; the at least one processing operator is configured to obtain one or more processing result by processing the initial multimodal data and send the processing result to the at least one output operator; the at least one output operator is configured to send output result generated based on the processing result to the multimodal data table. Detailed descriptions of the contents of processing multimodal data through the operator framework may be referred to relevant descriptions of FIG. 1 -FIG. 3 or FIG. 7.
[0173] In some embodiments, the processing module 620 may be configured that at least one read connector and at least one write connector, and the output result of the at least one output operator comprises target multimodal data; the at least one read connector is configured to obtain real data of the initial multimodal data from an actual storage address of the initial multimodal data and send the real data of the initial multimodal data to the at least one source operator; the at least one write connector is configured to write real data of the target multimodal data output by the at least one output operator to the actual storage address of the initial multimodal data. For more information about the read connector and the write connector, see FIG. 1 and its related description.
[0174] In some embodiments, the processing module 620 may be configured that receiving a first processing result from an upstream processing operator and processing the first processing result to generate a second processing result, and / or processing the initial multimodal data input by the source operator to generate the second processing result and sending the second processing result to a downstream processing operator. Detailed descriptions of the contents of the multilevel processing operators may be referred to relevant descriptions of FIG. 1-FIG. 3 and their related description.
[0175] FIG. 7 is a flow chart illustrating an exemplary process of processing data according to some embodiments of the present disclosure. Process 700 may be executed by the processor. The processor may establish a data connection with various components in the data processing system (e.g., the acquisition module and the processing module) by wired or wireless means in order to realize control of the data processing system. As shown in FIG. 7, process 700 includes the following operations.
[0176] In 710, a multimodal data table may be obtained.
[0177] In some embodiments, operation 710 may be performed by the acquisition module 610.
[0178] In some embodiments, the data in the multimodal data table comprises a combination of one or more of structured data, unstructured data of a single modality, or unstructured data of multimodalities.
[0179] Detailed descriptions of the contents of the multimodal data table, the structured data, and the unstructured data may be referred to relevant descriptions of FIG. 1.
[0180] The multimodal data tables may be constructed in a number of ways.
[0181] In some embodiments, the processor may obtain the multimodal data table based on the native data management system. The native data management system includes database management systems (e.g. Database Management System) , a data warehouses (e.g., a data warehouses implemented based on Hive) , or other database / data warehouse-like systems that support the management and use of the structured data in terms of table concepts, e.g., Presto, Trino, etc. For a more detailed description of the native data management system as well as the extensions, see FIG. 1 and its related description.
[0182] In some embodiments, the processor may also build the multimodal data table by designing data management system or a framework. For example, the processor may define the corresponding descriptive information or the metadata, etc., for the structured data and the unstructured data, respectively, which may be constructed to obtain the multimodal data table. For more information on designing a new data management system or framework, see FIG. 1 and its related description.
[0183] Through the setup of the multimodal data table, the unstructured modal data may be better managed, and the data modality of the stored unstructured data may be directly perceived, which facilitates the subsequent data analysis or data fusion.
[0184] In 720, the multimodal data table may be input into the operator framework for data processing to obtain processing result. In some embodiments, operation 720 may be performed by the processing module 620. In some embodiments, operation 720 may be performed by the operator framework as described elsewhere in the present disclosure. In some embodiments, the processing module 620 may be implemented by the operator framework as described elsewhere in the present disclosure.
[0185] More descriptions of the operator framework may be found elsewhere in the present disclosure (e.g., FIG. 1 -FIG. 3 and their related descriptions)
[0186] In some embodiments, the operator framework may perform data processing on the structured data in a multimodal data table, and the processing result is structured data. The processing is similar to traditional databases and data warehouses.
[0187] In some embodiments, the operator framework may perform data processing on the unstructured data in the multimodal data table, and the obtained processing result may includes: structured data, unstructured data of a single modality, unstructured data of multimodalities, structured data and a combination of the unstructured data of one or more modalities.
[0188] In some embodiments, the operator framework may perform data processing on both structured data and unstructured data in the multimodal data table, and the obtained processing result may includes: structured data, unstructured data of a single modality, unstructured data of multimodalities, structured data and a combination of the unstructured data of one or more modalities.
[0189] The processing result is a processed multimodal data table, and in some embodiments, the processing result may be displayed in an upper-tier application. The processing result are queried and used in the upper-tier application, including, but not limited to, displaying image content, playing audio / video, etc.
[0190] In some embodiments, the user may implement data processing for the various scenarios described above based on SQL, SQL-like languages, or drag and drop graphical.
[0191] See FIG. 1 -FIG. 3 and their related descriptions for more information on the processing result of the operator framework.
[0192] According to some embodiments of the present disclosure, data processing of multimodal data is enabled by providing an operator framework for analyzing and processing the multimodal data. Compared with conventional data processing methods, the operator framework may support the input of a plurality of different modal data and perform fusion analysis and processing of the plurality of different modal data, simplifying the steps of analyzing the multimodal data and realizing the common management of the structured data and the unstructured data of multimodalities, reducing the complexity and operation threshold of the multimodal data fusion analysis, and improving the efficiency of the multimodal data fusion analysis.
[0193] One or more embodiments of the present disclosure provide a computer-readable storage medium. The data processing method described in the above embodiments may be implemented based on a computer-readable storage medium.
[0194] Computer-readable storage media are described below in conjunction with FIG. 8. As shown in FIG. 8, a computer-readable storage medium 800 stores computer instructions (not shown in the figure) and program data 810, and when the computer reads the computer instructions (not shown in the figure) in the storage medium and / or processes the execution program data 810, the computer performs a data processing method comprising: obtaining a multimodal data table, inputting the multimodal data table into an operator framework for data processing to obtain a processing result. Detailed descriptions of the data processing method may be referred to relevant descriptions of FIG. 6.
[0195] Embodiments of the present application may be stored in a computer-readable storage medium when implemented as a software functional unit and sold or used as a stand-alone product. Based on this understanding, the technical solution of the present application may be embodied essentially or in part as a contribution to the prior art or in whole or in part in the form of a software product, which is a computer software product stored in a storage medium comprising instructions to cause a computer device (e.g., a personal computer, server, or network device, etc. ) or processor to perform all or some of the steps of the method described in the various embodiments of the present application. The aforementioned storage media include: USB flash drives, removable hard disks, read-only memories (ROM, Read-Only Memory) , random access memories (RAM, Random Access Memory) , disks or CD-ROMs, and other media that may store program code.
[0196] The basic concepts have been described above, and it is apparent to those skilled in the art that the foregoing detailed disclosure serves only as an example and does not constitute a limitation of the present disclosure. While not expressly stated herein, a person skilled in the art may make various modifications, improvements, and amendments to the present disclosure. Those types of modifications, improvements, and amendments are suggested in the present disclosure, so those types of modifications, improvements, and amendments remain within the spirit and scope of the exemplary embodiments of the present disclosure.
[0197] Also, the present disclosure uses specific words to describe embodiments of the present disclosure. such as "an embodiment" , "an embodiment" , and / or "some embodiments" means a feature, structure, or characteristic associated with at least one embodiment of the present disclosure. Accordingly, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" in different places in the present disclosure do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics of one or more embodiments of the present disclosure may be suitably combined.
[0198] Furthermore, unless expressly stated in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names as described in the present disclosure are not intended to qualify the order of the processes and methods of the present disclosure. While some embodiments of the invention that are currently considered useful are discussed in the foregoing disclosure by way of various examples, it is to be understood that such details serve only illustrative purposes, and that additional claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all amendments and equivalent combinations that are consistent with the substance and scope of the embodiments of the present disclosure. For example, although the implementation of various components described above may be embodied in a hardware device, it may also be implemented as a software only solution, e.g., an installation on an existing server or mobile device.
[0199] Similarly, it should be noted that in order to simplify the presentation of the disclosure of the present disclosure, and thereby aid in the understanding of one or more embodiments of the invention, the foregoing descriptions of embodiments of the present disclosure sometimes group multiple features together in a single embodiment, accompanying drawings, or in a description thereof. However, this method of disclosure does not imply that the objects of the present disclosure require more features than those mentioned in the claims. Rather, claimed subject matter may lie in less than all features of a single foregoing disclosed embodiment.
[0200] Some embodiments use numbers to describe the number of components, attributes, and it should be understood that such numbers used in the description of the embodiments are modified in some examples by the modifiers "about" , "approximately" , or "substantially" . ", "approximately" , or " generally" is used in some examples. Unless otherwise noted, the terms "about, " "approximately, " or "substantially" indicates that a ±20%variation in the stated number is allowed. Correspondingly, in some embodiments, the numerical parameters used in the present disclosure and claims are approximations, which may change depending on the desired characteristics of the individual embodiment. In some embodiments, the numerical parameters should take into account the specified number of valid digits and employ general place-keeping. While the numerical domains and parameters used to confirm the breadth of their ranges in some embodiments of the present disclosure are approximations, in specific embodiments such values are set to be as precise as possible within a feasible range.
[0201] For each patent, patent application, patent application disclosure, and other material cited in the present disclosure, such as articles, books, manuals, publications, documents, etc., the entire contents of which are hereby incorporated herein by reference. Application history documents that are inconsistent with or conflict with the contents of the present disclosure are excluded, as are documents (currently or hereafter appended to the present disclosure) that limit the broadest scope of the claims of the present disclosure. It should be noted that to the extent that the descriptions, definitions, and / or use of terms in the materials appurtenant to the present disclosure are inconsistent or in conflict with those set forth herein, the descriptions, definitions and / or use of terms in the present disclosure shall prevail.
[0202] Finally, it should be understood that the embodiments described herein are only used to illustrate the principles of the embodiments of the present disclosure. Other deformations may also fall within the scope of the present disclosure. As such, alternative configurations of embodiments of the present disclosure may be considered to be consistent with the teachings of the present disclosure as an example, not as a limitation. Correspondingly, the embodiments of the present disclosure are not limited to the embodiments expressly presented and described herein.
Claims
1.A multimodal operator framework wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator;the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data to the at least one processing operator;the at least one processing operator is configured to obtain one or more processing result by processing the initial multimodal data and send the processing result to the at least one output operator;the at least one output operator is configured to send one or more output result generated based on the processing result to the multimodal data table.2.The operator framework of claim 1, wherein the operator framework further comprises:at least one read connector and at least one write connector, and the output result of the at least one output operator comprises target multimodal data;the at least one read connector is configured to obtain real data of the initial multimodal data from an actual storage address of the initial multimodal data and send the real data of the initial multimodal data to the at least one source operator;the at least one write connector is configured to write real data of the target multimodal data output by the at least one output operator to the actual storage address of the initial multimodal data.3.The operator framework of claim 1, wherein the operator framework comprises multilevel processing operators, and for one of the multilevel processing operators:receiving a first processing result from an upstream processing operator and processing the first processing result to generate a second processing result, and / or processing the initial multimodal data input by the source operator to generate the second processing result and sending the second processing result to a downstream processing operator.4.The operator framework of any one of claims 1-3, wherein the operator framework further comprises:at least one association operator configured to associate output result from the source operator and / or the processing result from the at least one processing operator.5.The operator framework of claim 4, wherein an output of the at least one association operator is used as an input of the at least one output operator or as an input of the at least one processing operator.6.The operator framework of any one of claims 1-5, wherein the processing result comprise:at least one of data of another modalities converted from the data of each modality in the initial multimodal data, data of one or more specified modalities extracted from the initial multimodal data, or data of one or more specified modalities filtered from the initial multimodal data.7.The operator framework of any one of claims 1-6, wherein multimodal data table is configured to store metadata of the multimodal data, the metadata comprising at least one of a modal type field or an actual storage field;the modal type field is used to describe a file type of the multimodal data;the actual storage field is used to describe the actual storage address of real data of the multimodal data.8.A method of processing data, wherein the method comprises:obtaining at least one multimodal data table including initial multimodal data; andinputting the at least one multimodal data table into an operator framework as claimed in any one of claims 1-7 for data processing to obtain processing result.9.The method of claim 8, wherein the data in the multimodal data table comprises: at least one of structured data, unstructured data of a single modality, and unstructured data of multimodality.10.A data processing device wherein the device comprises at least one processor and at least one memory;the at least one memory for storing computer instructions;the at least one processor for executing at least a portion of the computer instructions to implement:obtaining at least one multimodal data table including initial multimodal data; andinputting the at least one multimodal data table into an operator framework for data processing to obtain processing result, wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator;the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data to the at least one processing operator;the at least one processing operator is configured to obtain one or more processing result by processing the initial multimodal data and send the processing result to the at least one output operator; andthe at least one output operator is configured to send one or more output result generated based on the processing result to the multimodal data table.11.The device of claim 10, wherein the operator framework further comprises:at least one read connector and at least one write connector, and the output of the at least one output operator comprises target multimodal data;the at least one read connector is configured to obtain real data of the initial multimodal data from an actual storage address of the initial multimodal data and send the real data of the initial multimodal data to the at least one source operator;the at least one write connector is configured to write real data of the target multimodal data output by the at least one output operator to the actual storage address of the initial multimodal data.12.The device of claim 10, wherein the operator framework comprises multilevel processing operators, and for of the multilevel processing operators:receiving a first processing result from an upstream processing operator and processing the first processing result to generate a second processing result, and / or processing the initial multimodal data input by the source operator to generate the second processing result and sending the second processing result to a downstream processing operator.13.The device of any one of claims 10-12, wherein the operator framework further comprises:at least one association operator configured to associate output result from the source operator and / or the processing result from the at least one processing operator.14.The device of claim 13, wherein an output of the at least one association operator is used as an input of the at least one output operator or as an input of the at least one processing operator.15.The device of any one of claims 10-14, wherein the processing result comprise:at least one of data of another modalities converted from the data of each modality in the initial multimodal data, data of one or more specified modalities extracted from the initial multimodal data, or data of one or more specified modalities filtered from the initial multimodal data.16.The device of any one of claims 10-15, wherein multimodal data table is configured to store metadata of the multimodal data, the metadata comprising at least one of a modal type field or an actual storage field;the modal type field is used to describe a file type of the multimodal data;the actual storage field is used to describe the actual storage address of real data of the multimodal data.17.A data processing system wherein the system comprises:an acquisition module obtaining at least one multimodal data table including initial multimodal data; anda processing module for inputting the at least one multimodal data tables into an operator framework for data processing to obtain processing result, wherein the operator framework comprises:at least one source operator, at least one processing operator, and at least one output operator;the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data into the at least one processing operator;the at least one processing operator is configured to obtain one or more processing result by processing the initial multimodal data and send the processing result to the at least one output operator; andthe at least one output operator is configured to send one or more output result generated based on the processing result to the multimodal data table.18.A computer-readable storage medium, the storage medium storing computer instructions, and when the computer reads the computer instructions in the storage medium, the computer performs a method of data processing comprising:obtaining at least one multimodal data table including initial multimodal data; andinputting the at least one multimodal data tables into an operator framework for data processing to obtain processing result, wherein the operator framework comprises: at least one source operator, at least one processing operator, and at least one output operator;the at least one source operator is configured to obtain initial multimodal data in a multimodal data table and input the initial multimodal data into the at least one processing operator;the at least one processing operator is configured to obtain one or more processing result by processing the initial multimodal data and send the processing result to the at least one output operator; andthe at least one output operator is configured to send one or more output result generated based on the processing result to the multimodal data table.
Citation Information
Patent Citations
Social relation analysis method and system based on multi-modal data and storage medium
CN115293920A
Multi-modal input digital character generation method and device, equipment and storage medium
CN117132864A
Operator framework, data processing method and device and computer storage medium
CN117931794A
VideoChat
US20210390138A1
System and Method for Multi-modality Soft-agent for Query Population and Information Mining
US20220114463A1