Communication method and device based on multi-modal data, equipment and storage medium
By pre-processing and standardizing multimodal data, the problems of inconsistency in data types and difficult semantic alignment in multimodal data processing are solved, efficient data interaction and intelligent processing are achieved, and the system's decision-making efficiency and interaction accuracy are improved.
Patent Information
- Application Number
- CN202510532371.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When processing multimodal data, traditional systems face problems such as inconsistent data types, large structural differences, and high semantic alignment difficulties, resulting in unsatisfactory information fusion effect, affecting the system's decision-making efficiency and interaction accuracy.
By acquiring multimodal data, pre-processing is performed and standardized based on the pre-constructed set of prompt words, and finally transmitted in a predetermined communication protocol, including modal recognition, structure-aware model, cross-modal semantic nesting and redundant content suppression and other technical means.
It improves the system's understanding and processing capabilities of complex information, significantly enhances the data interaction efficiency and intelligence level, and is suitable for cross-modal fusion and communication needs in multiple scenarios.
Smart Images

Figure CN120455540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a communication method, apparatus, device and storage medium based on multimodal data. Background Art
[0002] With the rapid development of artificial intelligence (AI), the demand for collecting and processing multimodal data is growing, particularly in applications such as smart terminals, autonomous driving, and human-computer interaction. Traditional systems often face challenges such as inconsistent data types, significant structural differences, and difficulty in semantic alignment. This leads to suboptimal information fusion, further impacting the system's decision-making efficiency and interaction accuracy. Therefore, a method that supports standardized processing and efficient transmission of multimodal data is urgently needed to improve data compatibility and intelligent processing capabilities, meeting the information interaction needs of complex application scenarios. Summary of the Invention
[0003] The present invention aims to provide a communication method, apparatus, device, and readable storage medium based on multimodal data to improve the above-mentioned problems. To achieve the above-mentioned objectives, the present invention adopts the following technical solutions:
[0004] In a first aspect, the present application provides a communication method based on multimodal data, comprising:
[0005] Acquiring multimodal data, wherein the multimodal data includes at least text, image, voice, and structured data;
[0006] Preprocessing the multimodal data to obtain preprocessed target data;
[0007] The target data is standardized based on a pre-constructed prompt word set to obtain revalued data;
[0008] The resetting data is communicated and transmitted based on a predetermined communication protocol.
[0009] In a second aspect, the present application further provides a communication device based on multimodal data, comprising:
[0010] an acquisition unit, configured to acquire multimodal data, wherein the multimodal data includes at least text, image, voice, and structured data;
[0011] a preprocessing unit, configured to preprocess the multimodal data to obtain preprocessed target data;
[0012] a first standardization unit, configured to perform standardization processing on the target data based on a pre-constructed prompt word set to obtain revalued data;
[0013] The transmission unit is configured to transmit the resetting data based on a predetermined communication protocol.
[0014] In a third aspect, the present application further provides a communication device based on multimodal data, comprising:
[0015] Memory for storing computer programs;
[0016] A processor is configured to implement the steps of the multimodal data-based communication method when executing the computer program.
[0017] In a fourth aspect, the present application further provides a readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned communication method based on multimodal data when executed by a processor.
[0018] The beneficial effects of the present invention are:
[0019] The present invention improves the system's ability to understand and process complex information through the fusion and standardized expression of multimodal data, and combines it with an efficient data transmission mechanism to significantly enhance the efficiency and intelligence level of data interaction, making it suitable for cross-modal fusion and communication needs in multiple scenarios.
[0020] Other features and advantages of the present invention will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 Schematic diagram of a multimodal data communication method according to an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of the structure of a communication device based on multimodal data according to an embodiment of the present invention;
[0024] Figure 3 Schematic diagram of the structure of a communication device based on multimodal data according to an embodiment of the present invention.
[0025] Markings in the figure: 10, acquisition unit; 20, preprocessing unit; 30, first standardization unit; 40, transmission unit; 800, communication device based on multimodal data; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0028] Example 1:
[0029] This embodiment provides a communication method based on multimodal data.
[0030] See also Figure 1 , the figure shows that the method includes step S10, step S20, step S30 and step S40.
[0031] Step S10. Acquire multimodal data, where the multimodal data includes at least text, image, voice, and structured data;
[0032] Specifically, different forms of raw data are collected in the target system or application environment to fully reflect the information status of the current scene. Multimodal data includes at least text data (such as user input and log information), image data (such as camera images), voice data (such as voice commands and environmental sounds), and structured data (such as sensor data and tabular data). The collection of multimodal data can be achieved through multi-source sensors, API interfaces, database calls, etc.
[0033] Step S20: preprocessing the multimodal data to obtain preprocessed target data;
[0034] Specifically, the acquired multimodal data is first cleaned and formatted to improve the availability and consistency of the data and ensure that subsequent standardized processing can be carried out efficiently.
[0035] Step S30: normalize the target data based on the pre-constructed prompt word set to obtain revalued data;
[0036] Specifically, based on a pre-built prompt word set, semantic recognition and reconstruction are performed on the target data to generate a standard expression data format and obtain the revalued data. The prompt word set includes at least data standard prompt words, supplementary semantic prompt words, enumeration prompt words, unit conversion prompt words, and fixed value rule prompt words.
[0037] Step S40: transmitting the reset value data based on a predetermined communication protocol;
[0038] Specifically, the standardized, revalued data is transmitted to the target system or module via a specific communication protocol, including but not limited to MQTT, HTTP, WebSocket, Modbus, and others, selected based on specific business needs. This transmission process ensures reliable data interaction and parsing consistency across multiple systems, supporting the subsequent processing, storage, and intelligent analysis of multimodal data.
[0039] It should be noted that step S20 includes steps S21 to S25:
[0040] Step S21. Performing modal recognition and data signal classification operations on the multimodal data to determine the modal type corresponding to each data unit, where the modal type includes but is not limited to text, image, voice, or table;
[0041] Specifically, considering that multimodal data often lacks a unified classification identifier, mixed processing of data from different modalities can easily lead to low processing efficiency and confusion in semantic understanding. Therefore, in this application, through modal recognition and signal classification mechanisms, it is possible to accurately distinguish different modal types such as text, images, voice, or tables, and to correspond different data processing strategies to different modal types, thereby effectively improving the system's data analysis accuracy and processing efficiency.
[0042] Step S22: parse the field structure of the multimodal data based on a preset structure-aware model, identify the hierarchical relationship between fields and the boundaries of semantic data blocks, and generate structure labels and field location information;
[0043] Specifically, the existing technology usually relies on static rules or template matching to perform structured parsing of unstructured or semi-structured data, such as identifying field locations and extracting key content through fixed formats. This approach is prone to failure when the format changes slightly and lacks the ability to adapt to complex hierarchies and diverse structures. Therefore, this application introduces a structure-aware model that can automatically identify field hierarchies and semantic boundaries, extract structural labels and field location information, and provide an accurate structural foundation for subsequent standardization.
[0044] Step S23: Using a cross-modal semantic nesting mechanism, align homologous fields or semantically related fields in different modalities to form an intermediate dataset after field alignment.
[0045] Specifically, in existing technologies, information between different modalities is often processed in isolation, resulting in fragmented context and insufficient semantic reconstruction capabilities. This application, by constructing a cross-modal semantic nesting mechanism, can effectively align homologous or semantically related fields in graphics, text, speech, and structured data, improving the depth and accuracy of information fusion and outputting a more consistent and expressive intermediate dataset.
[0046] Step S24: Redundant content is identified and suppressed on the intermediate data set, duplicate, conflicting, or ambiguous fields are deleted, and a set of fields with a high confidence level is retained to generate cleaned data content;
[0047] Specifically, this application uses a confidence scoring mechanism and an ambiguity resolution model to automatically identify and suppress redundant and conflicting fields in intermediate data sets, effectively retaining representative and credible fields and ensuring the high quality and consistency of data content.
[0048] Step S25. According to the preset field template or business structure rules, the cleaned data content is restructured and output as target data that conforms to the pre-processing standard format;
[0049] Specifically, by combining field templates with the business rule engine, the cleaned content can be flexibly reorganized, which not only achieves structural unification, but also has the advantages of strong business adaptability and good scalability, laying the foundation for subsequent data standardization and transmission.
[0050] It should be noted that step S22 includes steps S221 to S224:
[0051] Step S221: Input the multimodal data into a preset structural perception model, and the model outputs a perception recognition result including structural features such as field blocks, row and column logic, indentation levels, and adjacent field relationships;
[0052] Specifically, for example, by inputting a scanned image of a personal information form, the model can identify the structural features present therein, including: field blocks, such as "name", "date of birth", "contact information", etc.; row and column logic, where the name and date of birth are on the same row, and different rows represent different users; indentation levels, where if there are "province" or "street" under the "address" field, they are classified as subfields; and adjacent field relationships, where it is determined that "contact number" and "email address" are in the same semantic unit.
[0053] Step S222: Based on the perception recognition results, analyze the hierarchical structure relationship between fields in the multimodal data;
[0054] Specifically, based on the information output by the model, it is determined whether the field is a main field or a subfield, whether it belongs to the same group, etc., so as to establish a tree-like or hierarchical relationship between the fields.
[0055] For example, in a form containing address information: consider "Address" to be the parent field, and "Province", "City", and "Zip Code" to be child fields.
[0056] Step S223. Based on the structural information of the perception recognition result, determine the semantic consistency area of the multimodal data, identify the boundary position of each semantic data block, and generate a boundary recognition result;
[0057] Specifically, based on the structural information, it is determined which fields belong to the same semantic block (ie, have the same business logic or semantic target), and their start and end boundaries are identified for annotation.
[0058] Step S224. Based on the structural hierarchical relationship between fields and the boundary identification results, a structural label is assigned to each field, and corresponding field location information is generated;
[0059] Specifically, the system assigns labels to fields based on the structural hierarchy and boundary information, such as "main field", "subfield", "last-level field", etc., and records their location information in the original data (such as page number, coordinates, index, etc.).
[0060] It should be noted that step S30 includes steps S31 to S33:
[0061] Step S31. For each field in the target data, based on its corresponding modality type, structure tag information, and field location information, select and match the most relevant prompt word type and prompt word content from the prompt word set;
[0062] Specifically, for each field in the target data, the system first determines the modal type of the field (such as text, image, voice or structured data), and combines its structural label (such as "main field", "subfield") and positioning information (such as page number, position, context index) to retrieve and match the most relevant prompt word type and content in the prompt word set.
[0063] Step S32. Based on the matched prompt word type and field semantic results, the target data is standardized according to the preset standard rules;
[0064] Specifically, the system performs specific cleaning and conversion operations according to standardized rules based on the matched prompt word type and field semantics.
[0065] Step S33: Output the standardized target data as structured revalued data, which satisfies the data standards defined by the prompt word set in terms of field semantics, enumeration specifications, unit form, and value selection rules;
[0066] Specifically, the field data that has completed the standardization process will be structured and output to form revalued data that meets the predefined data standards. The revalued data needs to meet the requirements of clear semantic definition, consistent units, unified format, and compliant values.
[0067] It should be noted that step S31 includes steps S311 to S314:
[0068] Step S311. Perform field semantic analysis on each field in the target data in combination with its corresponding modality type, structure tag information, and field location information;
[0069] Specifically, the semantic analysis context is based on modality type (e.g., text, voice, image), structural labels (e.g., "header," "data unit"), and field location information (e.g., "row 2, column 3"). Natural language processing or semantic parsing models are then applied to identify the underlying meaning of the field. For example, if the field is "38.5," the modality type is "text," and it is located in the "body temperature" field cell of a table, its semantic meaning is initially determined to be "body temperature value."
[0070] Step S312: Perform semantic analysis and content extraction on each field, and build a field semantic model based on the semantic tags, contextual semantic relationships, and domain attributes analyzed for each field;
[0071] Specifically, the semantic labels of fields (such as "temperature value," "contact information," and "time expression") are extracted; context (such as the field context and the form module it belongs to) is combined with domain attributes (such as healthcare, finance, and transportation); and semantic embedding or knowledge graph representation is constructed to form a comparable semantic model. For example, the field "BP: 120 / 80" is labeled "blood pressure" and belongs to the "healthcare" domain. It coexists with other fields such as "body temperature" and "heart rate," further strengthening the semantics.
[0072] Step S313. Based on the field semantic model, determine the most likely applicable prompt word type for each field and determine a limited prompt word matching range;
[0073] Specifically, the semantic model determines which type of prompt word the field is most likely associated with, thereby eliminating irrelevant prompt word types and avoiding semantic conflicts or mismatches. For example, if the semantic model determines that the field is "blood pressure," the matching scope should be limited to "numeric format + unit conversion + normal range prompt words," and "gender enumeration prompt words" should not be matched.
[0074] Step S314. Match the prompt word content that is most relevant to the field semantics from the pre-constructed prompt word set within the limited prompt word matching range;
[0075] Specifically, within a set of limited prompt words, semantic similarity calculation or rule matching is performed. The prompt word with the highest similarity or the most logical fit is selected as the standardization basis for the field. For example, the field "38.5" matches the prompt word "Body temperature field: The value should be in ° C, ranging from 35 to 42." Based on this, the unit is added and verified to be reasonable.
[0076] It should be noted that step S32 includes steps S321 to S325:
[0077] Step S32. Based on the data standard prompt words, the field names and their corresponding values are uniformly named, formatted, and semantically aligned;
[0078] Specifically, data standard prompts are used to address issues such as non-standard field names, a mix of synonyms, and inconsistent data formats, ensuring standardized and consistent field semantics. For example, different names such as Order No., PO Number, and Order Number are uniformly recognized as "Order Number"; dates such as "2024 / 4 / 10, 2024.4.10" are converted to the uniform format "2024-04-10"; and "1,000 pieces" is converted to the numeric field "1000."
[0079] Step S32: Based on the supplementary semantic prompt words, semantically supplement and correct the semantically missing or ambiguous fields;
[0080] Specifically, when a field name is too short, its meaning is unclear, or the context lacks information, a prompt word can be used to complete its semantics and enhance the system's understanding. For example, when the field name is "Time," the semantics may be unclear. A semantic prompt word can be used to identify whether its contextual meaning is "Delivery Time" or "Payment Time." When the field name is "Material," the prompt word can be used to complete the description to "Cold-rolled Steel Plate - Material Type" after combining the context of "WLD001" as the value.
[0081] Step S32. Based on the enumeration value prompt, the non-standard enumeration class value is mapped to a unified enumeration standard;
[0082] Specifically, we map field values with enumeration characteristics (such as payment method, delivery status, and material type) to avoid confusion caused by different expressions. For example, we map "Cash on Delivery," "Cash Payment," and "Cash Delivery" to the standard enumeration item "Pay Now." We also map "Shipped," "Delivered," and "In Delivery" to the unified enumeration item "Shipped."
[0083] Step S32. Based on the unit conversion prompt word, realize the standard conversion and unified representation with the unit information field;
[0084] Specifically, we address the issue of inconsistent units and different expressions of physical quantities in field values, achieving numerical standardization and unified dimensional expression. For example, "2 tons" is converted to "2000kg"; "1000g" and "1 kilogram" are converted to "1kg"; and "3 boxes (10 pieces per box)" is converted to "30 pieces" based on the prompt word rules.
[0085] Step S32. Based on the value rule prompt word constraint field value range, correct, remove or mark the data that does not meet the rules;
[0086] Specifically, field values are validated and processed using pre-set rules (such as upper and lower numerical limits, reasonable time ranges, etc.) to ensure data validity. For example, if the date field value "2049-13-01" does not conform to the time format or exceeds the reasonable date range, the system will mark it as "invalid data." If the delivery quantity field value is "-50" or "0," the prompt word constraint "Must be a positive integer" will correct or eliminate it. If the "Payment Period" field value is "300 days" and exceeds the set upper limit (such as 180 days), a warning or red mark will be issued.
[0087] Example 2:
[0088] like Figure 2 As shown, this embodiment provides a communication device based on multimodal data, the device including:
[0089] An acquisition unit 10 is configured to acquire multimodal data, where the multimodal data includes at least text, image, voice, and structured data;
[0090] A preprocessing unit 20 is used to preprocess the multimodal data to obtain preprocessed target data;
[0091] A first standardization unit 30 is configured to perform standardization processing on the target data based on a pre-constructed prompt word set to obtain revalued data;
[0092] The transmission unit 40 is configured to transmit the resetting data based on a predetermined communication protocol.
[0093] In a specific embodiment disclosed in the present application, the pre-processing unit 20 includes:
[0094] A classification unit, configured to perform modal identification and data signal classification operations on multimodal data to determine the modal type corresponding to each data unit, including but not limited to text, image, voice, or table;
[0095] The parsing unit is used to parse the field structure of multimodal data based on a preset structure-aware model, identify the hierarchical relationship between fields and the boundaries of semantic data blocks, and generate structure labels and field location information;
[0096] The alignment unit is used to align homologous fields or semantically related fields in different modalities through a cross-modal semantic nesting mechanism to form an intermediate dataset after field alignment.
[0097] The suppression unit is used to identify and suppress redundant content in the intermediate data set, delete duplicate, conflicting or ambiguous fields, retain the field set with higher confidence, and generate cleaned data content;
[0098] The reorganization unit is used to restructure the cleaned data content according to the preset field template or business structure rules, and output the target data that conforms to the preprocessing standard format.
[0099] In a specific embodiment disclosed in this application, the first standardization unit 30 includes:
[0100] The first matching unit is configured to select and match the most relevant prompt word type and prompt word content from the prompt word set based on the corresponding modality type, structure tag information, and field location information of each field in the target data;
[0101] The second standardization unit is used to standardize the target data according to preset standard rules based on the matched prompt word type and field semantic results;
[0102] The output unit is used to output the standardized target data into structured revalued data, which meets the data standards defined by the prompt word set in terms of field semantics, enumeration specifications, unit form and value rules.
[0103] In a specific embodiment disclosed in this application, the first matching unit includes:
[0104] An analysis unit is used to perform field semantic analysis on each field in the target data in combination with its corresponding modality type, structure tag information, and field location information;
[0105] The extraction unit is used to perform semantic analysis and content extraction on each field, and build a field semantic model based on the semantic labels, contextual semantic relationships and domain attributes analyzed for each field;
[0106] A determination unit, configured to determine the most likely applicable prompt word type for each field based on the field semantic model, and to determine a limited prompt word matching range;
[0107] The second matching unit is configured to match the prompt word content most relevant to the field semantics from a pre-constructed prompt word set within a limited prompt word matching range.
[0108] In a specific embodiment disclosed in this application, the second standardization unit includes:
[0109] Naming unit, used to unify the naming, format conversion and semantic alignment of field names and their corresponding values based on data standard prompt words;
[0110] The supplementary unit is used to supplement and correct semantically missing or ambiguous fields based on supplementary semantic clue words;
[0111] A mapping unit, used to map non-standard enumeration class values to a unified enumeration standard based on enumeration value prompt words;
[0112] A conversion unit, used to implement standard conversion and unified representation with a unit information field based on a unit conversion prompt word;
[0113] The fixed value unit is used to constrain the field value range based on the fixed value rule prompt word and correct, eliminate or mark the data that does not meet the rules.
[0114] In a specific embodiment disclosed in this application, the parsing unit includes:
[0115] The input unit is used to input multimodal data into a preset structural perception model. The model outputs the perception recognition results containing structural features such as field blocks, row and column logic, indentation levels, and adjacent field relationships.
[0116] The recognition unit is used to analyze the structural hierarchical relationship between fields in the multimodal data based on the perception recognition results;
[0117] A determination unit, configured to determine the semantic consistency region of the multimodal data based on the structural information of the perception recognition result, identify the boundary position of each semantic data block, and generate a boundary recognition result;
[0118] The generation unit is used to assign a structural label to each field based on the structural hierarchical relationship and boundary recognition results between the fields, and generate corresponding field positioning information.
[0119] It should be noted that, regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0120] Example 3:
[0121] Corresponding to the above method embodiment, this embodiment also provides a communication device based on multimodal data. The communication device based on multimodal data described below and the communication method based on multimodal data described above can refer to each other.
[0122] Figure 3 FIG is a block diagram of a communication device 800 based on multimodal data according to an exemplary embodiment. Figure 3 As shown, the multimodal data-based communication device 800 may include: a processor 801 and a memory 802. The multimodal data-based communication device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0123] The processor 801 is used to control the overall operation of the multimodal data-based communication device 800 to complete all or part of the steps in the multimodal data-based communication method described above. The memory 802 is used to store various types of data to support the operation of the multimodal data-based communication device 800. This data may include, for example, instructions for any application or method operating on the multimodal data-based communication device 800, as well as application-related data, such as contact data, sent and received messages, pictures, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the communication device 800 based on multimodal data and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module, an NFC module.
[0124] In an exemplary embodiment, the multimodal data-based communication device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-mentioned multimodal data-based communication method.
[0125] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the aforementioned multimodal data-based communication method. For example, the computer-readable storage medium may be the aforementioned memory 802 including the program instructions. The program instructions may be executed by the processor 801 of the multimodal data-based communication device 800 to implement the aforementioned multimodal data-based communication method.
[0126] Example 4:
[0127] Corresponding to the above method embodiment, this embodiment further provides a readable storage medium. The readable storage medium described below and the communication method based on multimodal data described above can refer to each other.
[0128] A readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the multimodal data-based communication method of the above-mentioned method embodiment.
[0129] The readable storage medium may specifically be any readable storage medium that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0130] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0131] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A communication method based on multimodal data, characterized in that: include: Acquiring multimodal data, wherein the multimodal data includes at least text, image, voice, and structured data; Preprocessing the multimodal data to obtain preprocessed target data; Standardizing the target data based on a pre-constructed prompt word set to obtain revalued data; The resetting data is communicated and transmitted based on a predetermined communication protocol.
2. The communication method based on multimodal data according to claim 1, characterized in that , preprocessing the multimodal data to obtain preprocessed target data, including: Performing modality recognition and data signal classification operations on the multimodal data to determine the modality type corresponding to each data unit, the modality type including but not limited to text, image, voice, or table; Performing field structure analysis on the multimodal data based on a preset structure-aware model, identifying hierarchical relationships between fields and semantic data block boundaries, and generating structure labels and field location information; Through the cross-modal semantic nesting mechanism, homologous fields or semantically related fields in different modalities are aligned to form an intermediate dataset after field alignment; Identifying and suppressing redundant content on the intermediate data set, deleting duplicate, conflicting, or ambiguous fields, retaining a set of fields with a high confidence level, and generating cleaned data content; According to the preset field templates or business structure rules, the cleaned data content is restructured and reorganized to output the target data that conforms to the preprocessing standard format.
3. The communication method based on multimodal data according to claim 2, characterized in that The target data is standardized based on a pre-constructed prompt word set to obtain revalued data. The prompt word set includes data standard prompt words, supplementary semantic prompt words, enumeration value prompt words, unit conversion prompt words and value rule prompt words, including: For each field in the target data, combined with its corresponding modality type, structure tag information, and field location information, the most relevant prompt word type and prompt word content are selected and matched from the prompt word set; Based on the matched prompt word type and field semantic results, the target data is standardized according to preset standard rules; The standardized target data is output as structured revalued data, which meets the data standards defined by the prompt word set in terms of field semantics, enumeration specifications, unit form and value rules.
4. The communication method based on multimodal data according to claim 3, characterized in that ,For each field information in the target data, combined with its corresponding modality type, structure label information and field positioning information, the most relevant prompt word type and prompt word content are selected and matched from the prompt word set, including: For each field in the target data, perform field semantic analysis based on its corresponding modality type, structure tag information, and field location information; Perform semantic analysis and content extraction on each field, and build a field semantic model based on the semantic labels, contextual semantic relationships, and domain attributes analyzed for each field; Based on the field semantic model, determine the most likely applicable prompt word type for each field, and determine a limited prompt word matching range; Within the limited prompt word matching range, the prompt word content most relevant to the field semantics is matched from the pre-constructed prompt word set.
5. A communication device based on multimodal data, characterized in that: include: an acquisition unit, configured to acquire multimodal data, wherein the multimodal data includes at least text, image, voice, and structured data; a preprocessing unit, configured to preprocess the multimodal data to obtain preprocessed target data; a first standardization unit, configured to perform standardization processing on the target data based on a pre-constructed prompt word set to obtain revalued data; The transmission unit is configured to transmit the resetting data based on a predetermined communication protocol.
6. The multimodal data-based communication device according to claim 5, characterized in that: The pre-processing unit comprises: a classification unit, configured to perform modality recognition and data signal classification operations on the multimodal data to determine a modality type corresponding to each data unit, wherein the modality type includes but is not limited to text, image, voice, or table; A parsing unit, configured to perform field structure parsing on the multimodal data based on a preset structure-aware model, identify hierarchical relationships between fields and semantic data block boundaries, and generate structure labels and field location information; The alignment unit is used to align homologous fields or semantically related fields in different modalities through a cross-modal semantic nesting mechanism to form an intermediate dataset after field alignment. a suppression unit, configured to identify and suppress redundant content in the intermediate data set, delete duplicate, conflicting, or ambiguous fields, retain a set of fields with a higher confidence level, and generate cleaned data content; The reorganization unit is used to restructure the cleaned data content according to the preset field template or business structure rules, and output the target data that conforms to the preprocessing standard format.
7. The multimodal data-based communication device according to claim 6, characterized in that: The first standardization unit includes: The first matching unit is configured to select and match the most relevant prompt word type and prompt word content from the prompt word set based on the corresponding modality type, structure tag information, and field location information of each field in the target data; A second standardization unit is configured to perform standardization processing on the target data according to preset standard rules based on the matched prompt word type and field semantic results; The output unit is used to output the standardized target data as structured revalued data, wherein the revalued data meets the data standards defined by the prompt word set in terms of field semantics, enumeration specifications, unit form and value rules.
8. The multimodal data-based communication device according to claim 7, characterized in that: The first matching unit includes: An analysis unit is used to perform field semantic analysis on each field in the target data in combination with its corresponding modality type, structure tag information, and field location information; The extraction unit is used to perform semantic analysis and content extraction on each field, and build a field semantic model based on the semantic labels, contextual semantic relationships and domain attributes analyzed for each field; a determination unit, configured to determine the most likely applicable prompt word type for each field based on the field semantic model, and determine a limited prompt word matching range; The second matching unit is configured to match the prompt word content most relevant to the field semantics from a pre-constructed prompt word set within the limited prompt word matching range.
9. A communication device based on multimodal data, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the multimodal data-based communication method according to any one of claims 1 to 4 when executing the computer program.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the communication method based on multimodal data according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Multi-modal data fusion method based on compound collaborative structure feature recombination network
CN113378989A
Multi-modal data timestamp synchronization method and device, electronic equipment and storage medium
CN117294376A
Text recognition method and device and text recognition model training method and device
CN117351507A
Multi-modal semantic communication method, system and equipment based on large model and medium
CN118350416A
Clinical data entry method and device, electronic equipment and storage medium
CN119274760A
Cited By
Universal scene retrieval analysis method and system based on multi-modal feature fusion
CN120821872A
A general scene retrieval analysis method and system based on multi-modal feature fusion
CN120821872B