Research and development document processing method and device
Through the multimodal document analysis engine and deep learning model, the problems of multimodal information integration and intelligent identification in the existing technology are solved, and efficient management of R&D documents and accumulation of knowledge assets are realized.
Patent Information
- Application Number
- CN202510249868.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing R&D document processing technology is difficult to effectively integrate multimodal information such as text, voice, and images, resulting in incomplete recording of the R&D process, poor information relevance, and lack of intelligent recognition and semantic understanding capabilities.
The multimodal document analysis engine is used to identify document formats, speech recognition, image recognition and content extraction to generate original data; through cross-modal preprocessing and semantic alignment networks, different types of data are converted into standardized feature representations; deep learning models are used to identify R&D elements and establish multimodal relationship maps to realize intelligent extraction and correlation analysis of information.
It effectively solves the shortcomings of multimodal information fusion and knowledge system construction, provides strong support for R&D process management and knowledge asset accumulation, and significantly improves the standardized management level of R&D documents.
Smart Images

Figure CN120087351A_ABST
Abstract
Description
Technical Field This application relates to the field of data processing, and in particular, to a method and device for processing R & D documents. Background Art There are obvious deficiencies in the existing R & D document processing technologies. Traditional methods mainly target the processing of single-type documents and cannot effectively integrate multi-modal information such as text, speech, and images, resulting in incomplete records of the R & D process and poor information relevance. At the same time, the existing methods lack the intelligent recognition and semantic understanding capabilities for R & D elements.
[0001] In addition, there are bottlenecks in cross-modal information processing and alignment in the existing technologies. Most systems fail to achieve effective fusion and standardized processing of different modal data, making it difficult to construct a unified feature representation space. The process of document sorting and archiving lacks systematicness and standardization, affecting the effective inheritance and utilization of R & D achievements.
[0002] There are technical shortcomings in multi-modal semantic association and knowledge graph construction in the existing systems. The deep relationships between R & D elements are lacking in excavation, and a complete R & D knowledge system cannot be formed. Solving these problems is of great significance for improving the efficiency of R & D management and the value of knowledge assets. Summary of the Invention In view of the problems in the existing technologies, this application provides a method and device for processing R & D documents, which can effectively solve the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provide strong support for R & D process management and knowledge asset accumulation, and significantly improve the standardized management level of R & D documents.
[0003] To solve at least one of the above problems, this application provides the following technical solutions: In a first aspect, this application provides a method for processing R & D documents, including: Establishing a multi-modal document parsing engine, where the multi-modal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. Inputting the R & D document into the document format recognition module for format type recognition, inputting the R & D process recording into the speech recognition module for speech-to-text processing, inputting the experimental pictures and equipment photos into the image recognition module for feature extraction and scene recognition, and using the content extraction module to extract the text content, speech text, and image annotations to generate raw data containing multi-modal information; Perform cross-modal preprocessing on the original data. Divide the original data into text blocks, speech blocks, and image blocks according to the information type. Perform word segmentation and part-of-speech tagging on the text blocks, construct text semantic vectors. Perform speaker recognition and speech segmentation on the speech blocks, extract speech feature vectors. Perform object detection and scene classification on the image blocks, generate image feature vectors. Input the text semantic vectors, speech feature vectors, and image feature vectors into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space; Input the standardized document data into a multi-modal content understanding model trained based on deep learning. Identify the R & D elements in the document through the multi-modal content understanding model, including R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D results. Conduct cross-modal semantic association analysis on the identified R & D elements, establish a multi-modal relationship map containing text descriptions, speech records, and image evidence, and reorganize the standardized document data according to a unified document template format based on the multi-modal relationship map to generate R & D document data integrating multi-modal information.
[0004] Furthermore, it includes: Construct a document format recognition network based on a neural network, extract document feature vectors from the document binary stream, perform classification training on the document feature vectors using a convolutional neural network to obtain a document format classification model, identify the format type of the input document based on the document format classification model, select the corresponding parsing rule library according to the recognition result to parse the document structure, and associate and save the parsed document structure information with the original document data; Construct a speech feature extraction network and an image feature extraction network. Use Mel-frequency cepstral coefficients to extract acoustic features from the R & D process recordings, input the acoustic features into a speech recognition model trained based on a long short-term memory network for text conversion. Perform convolutional feature extraction and residual network training on experimental pictures and equipment photos to obtain a scene classification model, classify and label the picture content through the scene classification model, and integrate the converted speech text and picture annotation information with the document structure information to obtain multi-modal document data.
[0005] Furthermore, it includes: Establish a document segmentation processing engine, segment the document content based on the document structure information, construct a text classification model using a bidirectional long short-term memory network, classify the theme of the segmented text fragments, identify the theme category of each text fragment, segment the speech text according to the timestamp information, calculate the semantic similarity of the segmented speech text fragments, and merge the speech text fragments with a semantic similarity higher than the preset threshold to generate speech text paragraphs; Construct a cross-modal feature fusion network to semantically match the topic information of text segments with the speech text paragraphs, extract keywords from the image annotation information, calculate the semantic correlation degree between the text segments, speech paragraphs and image annotations based on the attention mechanism, and use the multi-head attention mechanism to perform weighted fusion on the semantic correlation degree to generate a fused multi-modal document feature matrix. Combine the multi-modal document feature matrix with the document structure information to form the original data containing multi-modal information.
[0006] Further, it includes: Construct a block division model, establish information type recognition rules based on the structure marking information in the original data, perform sequence annotation on the original data using the conditional random field algorithm, identify the boundary positions of text, speech, and images, divide the original data into independent blocks according to the boundary positions, establish a text tokenization model based on word vectors, perform tokenization processing on the text blocks, use the hidden Markov model for part-of-speech tagging, and input the tagging results into a pre-trained text encoder to generate text semantic vectors. Construct a multi-modal feature extraction network, use the voiceprint recognition algorithm to extract the speaker's acoustic features from the speech blocks, train a speaker recognition model based on the acoustic features, segment the speech signal according to the silent segments, extract the acoustic feature parameters of the segmented speech segments, convert the acoustic feature parameters into speech feature vectors through a recurrent neural network encoder, apply an object detection network to the image blocks to identify the key objects in the images, use a deep residual network to classify the scenes of the images, and fuse the object detection results and the scene classification results to generate image feature vectors.
[0007] Further, it includes: Construct a cross-modal feature mapping network, train a feature transformation model using the adversarial learning method, input the text semantic vectors, speech feature vectors, and image feature vectors into the corresponding feature transformation models respectively to generate feature representations with the same dimension, calculate the mutual information loss for the transformed feature representations, optimize the parameters of the feature transformation model through backpropagation, so that different modal features have similar distribution characteristics in the unified feature space, and calculate the semantic similarity matrix between different modal features based on the cosine similarity. Construct a feature alignment optimization network, establish feature alignment constraints based on the semantic similarity matrix, perform relational reasoning on the feature representations using a graph neural network, perform attention-weighted fusion on the reasoning results and the feature representations, perform dimensionality reduction processing on the fused features through a multi-layer perceptron, normalize the dimensionality-reduced features to obtain a normalized representation, and establish a mapping relationship between the normalized representations with corresponding relationships and the original features to generate the normalized document data containing multi-modal alignment information.
[0008] Further, it includes: Construct a multi-modal feature sequence encoding network. Use the position encoding method to encode the position information of the feature sequence in the normalized document data. Input the encoded feature sequence into the encoder based on the Transformer architecture to perform multi-head self-attention calculation on the feature sequence to obtain the context representation. Use a recurrent neural network to perform temporal modeling on the context representation. Input the modeling result into a conditional random field for sequence labeling, and classify and label the document content based on a preset R & D element annotation system to identify the text segments corresponding to the R & D objectives, technical solutions, innovation points, R & D processes, experimental data, and R & D results; Construct an element recognition optimization network. Train an element classifier based on a pre-annotated R & D document dataset. Input the sequence labeling result into the element classifier for multi-label classification. Calculate the cross-entropy loss between the classification result and the preset label. Use the gradient descent method to optimize the classifier parameters. Re-extract features and perform classification prediction on the text segments with a classification confidence lower than the threshold. Integrate the optimized classification result with the sequence labeling result to generate the final R & D element recognition result.
[0009] Furthermore, it includes: Construct a multi-modal relationship reasoning network. Use semantic similarity calculation for the identified R & D elements to establish an association matrix between elements. Based on the graph attention network, sparsify the association matrix to obtain the initial graph structure. Use the feature representations of different modalities as the attribute information of the graph nodes. Use the message passing mechanism to perform feature aggregation on the graph structure, update the semantic representations of the nodes through a multi-layer graph convolutional network, calculate the relationship strength between nodes based on the updated semantic representations, and construct a multi-modal relationship graph containing text descriptions, voice records, and image evidence; Construct a document recombination and generation network. Establish document structure constraints based on a preset document template. Map the nodes in the multi-modal relationship graph according to the template structure. Use a graph-to-sequence generation model to convert the graph structure into a text sequence. Optimize the coherence at the discourse level for the generated text sequence. Insert the corresponding voice records and image evidence into the corresponding positions according to the semantic association relationship. Use a text polishing model to optimize the language of the inserted document content to generate R & D document data integrating multi-modal information.
[0010] In the second aspect, the present application provides a device for processing R & D documents, including: A multi-modal analysis module for establishing a multi-modal document analysis engine. The multi-modal document analysis engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. The R & D document is input into the document format recognition module for format type recognition, the R & D process recording is input into the speech recognition module for speech-to-text processing, the experimental pictures and equipment photos are input into the image recognition module for feature extraction and scene recognition, and the content extraction module is used to extract the text content, speech text, and image annotations to generate raw data containing multi-modal information; A cross-modal processing module for performing cross-modal preprocessing on the raw data. The raw data is divided into text blocks, speech blocks, and image blocks according to the information type. The text blocks are tokenized and part-of-speech tagged to construct text semantic vectors, the speech blocks are speaker-identified and speech-segmented to extract speech feature vectors, the image blocks are object-detected and scene-classified to generate image feature vectors, and the text semantic vectors, speech feature vectors, and image feature vectors are input into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space; A document standardization module for inputting the standardized document data into a multi-modal content understanding model trained based on deep learning. The R & D elements in the document are identified through the multi-modal content understanding model, including R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D results. Cross-modal semantic association analysis is performed on the identified R & D elements to establish a multi-modal relationship graph containing text descriptions, speech records, and image evidence, and the standardized document data is reorganized according to a unified document template format based on the multi-modal relationship graph to generate R & D document data integrating multi-modal information.
[0011] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the R & D document processing method described above are implemented.
[0012] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the R & D document processing method described above are implemented.
[0013] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the R & D document processing method described above are implemented.
[0014] As can be seen from the above technical solutions, the present application provides a method and device for processing R & D documents, which can effectively solve the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provide strong support for R & D process management and knowledge asset accumulation, and significantly improve the standardized management level of R & D documents. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a schematic flowchart of the method for processing R & D documents in the embodiments of the present application; Figure 2 It is a structural diagram of the device for processing R & D documents in the embodiments of the present application; Figure 3 It is a schematic structural diagram of the electronic device in the embodiments of the present application.
[0016] Reference Signs: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION OF THE EMBODIMENTS To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0017] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0018] In view of the problems existing in the prior art, the present application provides a research and development document processing method and device. By constructing a multi-modal document parsing engine, unified processing of text, speech, and image information is achieved. Through cross-modal preprocessing and semantic alignment network, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify research and development elements, establishes a multi-modal relationship graph, and realizes intelligent extraction and correlation analysis of various types of information in research and development documents. This method effectively solves the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provides strong support for research and development process management and knowledge asset accumulation, and significantly improves the standardized management level of research and development documents.
[0019] In order to effectively solve the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provide strong support for research and development process management and knowledge asset accumulation, and significantly improve the standardized management level of research and development documents, an embodiment of a research and development document processing method is provided in the present application. Refer to Figure 1 , the research and development document processing method specifically includes the following content: Step S101: Establish a multi-modal document parsing engine. The multi-modal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. Input the research and development document into the document format recognition module for format type recognition, input the research and development process recording into the speech recognition module for speech-to-text processing, input the experimental pictures and equipment photos into the image recognition module for feature extraction and scene recognition, and use the content extraction module to extract the text content, speech text, and image annotations to generate raw data containing multi-modal information; Optionally, in this embodiment, for the research and development document processing scenario, a multi-modal document parsing engine based on deep learning is first constructed. Considering that research and development documents usually include various formats such as Word documents, PDF files, charts, and experimental records, the document format recognition module adopts a deep convolutional neural network architecture based on VGG16. This network extracts the visual features of the document through five convolutional blocks. Each convolutional block contains 2-3 convolutional layers and a max pooling layer. The convolutional kernel size is 3×3, and the pooling window is 2×2. The fully connected layer of the network maps the extracted features to a predefined document category space and outputs the confidence scores of various formats.
[0020] In this embodiment, when processing the recordings of the R & D process, the speech recognition module adopts an improved end-to-end speech recognition architecture. First, logarithmic Mel spectrogram features are used to extract acoustic features, with the window length set to 25 ms and the step size to 10 ms. The acoustic sequence after feature extraction is input into a multi-layer bidirectional LSTM network, which contains 4 layers of LSTM, and the number of hidden units in each layer is 512. On top of the LSTM, a multi-head attention mechanism based on positional encoding is designed, with the number of heads set to 8, to capture the long-range dependencies of the acoustic features. The decoder uses a CTC-based acoustic model and integrates a Transformer-based language model, which can accurately recognize the professional terms and technical nouns in the R & D process.
[0021] In this embodiment, the image recognition module designs a two-stage recognition process according to the characteristics of experimental pictures and equipment photos. In the first stage, the YOLOv5 object detection network is used to identify key objects such as experimental equipment, instruments, and experimental materials in the image. The backbone feature extraction network of the network adopts the CSPDarknet53 structure, and the spatial pyramid pooling is used to enhance the detection ability for multi-scale targets. In the second stage, ResNet50 is used for scene classification. This network effectively alleviates the problem of gradient disappearance in deep networks through residual connections and can accurately distinguish different scene types such as laboratories and engineering sites.
[0022] The content extraction module of this embodiment adopts a hierarchical processing strategy. For text content, first, a rule-based layout analysis method is used to divide the document structure, including titles, texts, chart descriptions, etc. Then, the BERT pre-trained model is applied for semantic understanding. The input of the model includes the text sequence and positional encoding, and the context-related semantic features are extracted through 12 layers of Transformer encoders. The model is fine-tuned using domain-specific corpora, which improves the understanding ability of professional terms in the R & D field.
[0023] In this embodiment, when processing speech texts, a speaker segmentation and role recognition mechanism is innovatively introduced. The i-vector-based speaker recognition technology is adopted to extract 600-dimensional voiceprint feature vectors, and the speaker identity is judged through PLDA scoring. This enables the system to distinguish the speech contents of different roles such as R & D personnel and review experts, providing important context information for subsequent semantic analysis. For image annotation, an instance segmentation network based on Faster R-CNN is adopted to generate pixel-level object masks, and GCN is combined for relationship reasoning to identify the spatial and functional associations between objects in the image.
[0024] This embodiment designs a multi-modal feature fusion strategy, which uses an attention mechanism to dynamically allocate weights to features of different modalities. When processing experimental data, the system increases the weight of numerical content according to the context; when describing experimental phenomena, it enhances the importance of image features. Through this adaptive weight adjustment, the reasonable fusion of different types of information is ensured. The feature fusion adopts a concat-attention structure. First, the features of different modalities are concatenated, and then the interaction relationships between modalities are learned through a self-attention mechanism.
[0025] In the original data generation link of this embodiment, the temporal alignment of multi-modal information is realized. The system analyzes the time stamps in the document, speech timestamps, and picture EXIF information to establish the temporal correlation of different modal data. The dynamic time warping algorithm is used to process the time scale differences between different modalities to ensure the temporal consistency of the generated original data. At the same time, a consistency check mechanism based on a knowledge graph is introduced to discover and correct semantic conflicts between multi-modal information.
[0026] Through the above technological innovations, this embodiment effectively solves problems such as difficult parsing of multi-source heterogeneous data, incomplete information extraction, and insufficient modal fusion in traditional R & D document processing. In practical applications, this solution can accurately identify the format types of various R & D documents, complete speech transcription with high quality, accurately extract image content, and achieve the effective integration of multi-modal information. It is particularly suitable for document processing in large-scale R & D projects, significantly improving the automation level and accuracy of document processing. The modular design of this solution also facilitates optimization and expansion according to the characteristics of different R & D fields.
[0027] Step S102: Perform cross-modal preprocessing on the original data. Divide the original data into text blocks, speech blocks, and image blocks according to the information type. Perform word segmentation and part-of-speech tagging on the text blocks to construct text semantic vectors, perform speaker recognition and speech segmentation on the speech blocks to extract speech feature vectors, perform object detection and scene classification on the image blocks to generate image feature vectors, and input the text semantic vectors, speech feature vectors, and image feature vectors into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space; Optionally, this embodiment designs a complete cross-modal preprocessing process for multi-modal data in R & D documents. First, based on document structure tags and content features, the hierarchical clustering algorithm is used to divide the original data into different information blocks. The clustering features include content type identifiers, location information, time stamps, etc., and the optimal block boundaries are determined through a dynamic programming algorithm. The system can accurately identify different types of information units such as text paragraphs, meeting recording segments, and experimental process photos in R & D reports.
[0028] In this embodiment, when processing text blocks, an improved word segmentation model based on deep learning is adopted. This model takes character-level features as input and learns the context dependencies between characters through a multi-layer bidirectional LSTM network. To improve the recognition accuracy of professional terms, the model integrates a domain dictionary and weights the dictionary matching results through an attention mechanism. Part-of-speech tagging uses a sequence tagging model based on BERT, which is pre-trained on domain-specific corpora and can accurately tag special parts of speech such as proper nouns and technical indicators in R & D documents.
[0029] This embodiment innovatively constructs hierarchical text semantic vectors. First, the Word2Vec model is used to learn word-level semantic representations. The model adopts the Skip-gram architecture with a window size of 5 and outputs 300-dimensional word vectors. Then, a hierarchical attention network is used to construct sentence-level semantic vectors. This network contains two levels: word-level attention and sentence-level attention, and can capture important information within sentences and logical relationships between sentences. Finally, the document structure information is integrated through a graph convolutional network to generate a text semantic representation that takes into account the global context.
[0030] For speech blocks, this embodiment implements a speaker recognition system based on deep features. First, 60-dimensional MFCC features and their first- and second-order differences are extracted to construct an acoustic feature sequence. Then, a voiceprint feature is extracted through a residual network, which contains 5 residual blocks, and each residual block contains two convolutional layers and a shortcut connection. To improve the robustness of speaker recognition, an adversarial training strategy is introduced to enhance the generalization ability of the model by adding Gaussian noise. Voice segmentation uses a dual-threshold detection algorithm based on VAD and optimizes the position of the segmentation point by combining prosodic features.
[0031] In the speech feature vector extraction stage of this embodiment, a multi-scale feature fusion mechanism is designed. First, local acoustic features are extracted through a convolutional neural network, then the self-attention mechanism is used to capture long-range dependencies, and finally, different-scale features are selectively fused through a gating mechanism. The output dimension of the feature extraction network is 512, which contains information in multiple aspects such as speaker identity features, emotional features, and semantic features. The system also implements a speaker role annotation function, which can distinguish different roles such as the host, the presenter, and the review expert.
[0032] In this embodiment, a two-stage processing strategy is adopted for image blocks. First, an improved Faster R-CNN network is used for object detection. The backbone network adopts the ResNet101 structure, and multi-scale feature maps are constructed through FPN. The network is fine-tuned on object detection datasets such as experimental equipment and instrumentation, improving the recognition ability of professional equipment. Scene classification adopts the DenseNet architecture, and the feature utilization efficiency is improved through dense connections. The network outputs a 2048-dimensional feature vector, which contains rich visual information such as object categories, spatial layouts, and scene attributes.
[0033] In this embodiment, a cross-modal alignment network is innovatively designed. The network adopts a dual learning framework, which includes three independent encoders that process text, speech, and image features respectively. The encoders adopt similar network structures, including multiple layers of Transformer blocks, and extract intra-modal relational features through self-attention mechanisms. In the training stage, a triplet loss function is introduced to guide feature alignment. For cross-modal sample pairs that are semantically related, the distance between them in the feature space is minimized; for unrelated sample pairs, the feature distance is ensured to be greater than a preset threshold.
[0034] In this embodiment, the unification of the feature space based on adversarial learning is realized. A discriminator network is connected after the encoder of each modality. The discriminator attempts to distinguish the feature distributions of different modalities, while the encoder makes the feature distributions of different modalities tend to be consistent through adversarial training. This method effectively eliminates modality differences, enabling different types of information to be compared and fused in a unified feature space.
[0035] Through the above technological innovations, this embodiment effectively solves the problems of multi-modal data preprocessing and alignment in R & D documents. This solution can accurately process text content in professional fields, accurately identify speakers in multi-person meeting scenarios, effectively extract image information during the experimental process, and achieve semantic-level alignment of different modal data. In practical applications, it significantly improves the accuracy of subsequent content understanding and knowledge extraction. The modular design of the solution also supports optimization and expansion according to the characteristics of different R & D fields.
[0036] Step S103: Input the standardized document data into a multi-modal content understanding model trained based on deep learning. Through the multi-modal content understanding model, identify the R & D elements in the document, including R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D results. Conduct cross-modal semantic association analysis on the identified R & D elements, establish a multi-modal relationship graph containing text descriptions, voice records, and image evidence, and reorganize the standardized document data according to a unified document template format based on the multi-modal relationship graph to generate R & D document data integrating multi-modal information.
[0037] Optionally, in this embodiment, a dedicated multi-modal content understanding model is first designed according to the characteristics of R & D documents. This model adopts a hierarchical Transformer architecture, including three key levels: the underlying feature encoding layer, the middle-level semantic understanding layer, and the high-level knowledge reasoning layer. In the underlying feature encoding, BERT is used as the encoder for text, Conformer is used to encode speech features, and Swin Transformer is used to process image features. The features of each modality are first preprocessed in their respective encoders to generate modality-specific context representations.
[0038] In this embodiment, a cross-modal attention mechanism is innovatively introduced in the middle-level semantic understanding layer. This mechanism allows information interaction between different modalities. For example, when processing experimental data, the model can simultaneously focus on numerical descriptions in the text, experimental instructions in the speech, and instrument readings in the image. The attention weights are calculated through the softmax function, where Qi, Kj, and Vj represent the query vector, key vector, and value vector respectively, and i and j represent different modalities: Attention(Qi, Kj, Vj) = softmax(QiKj^T / √d)Vj, where d is the feature dimension. This mechanism enables the model to accurately capture the semantic associations between cross-modalities.
[0039] In the high-level knowledge reasoning layer of this embodiment, a R & D element recognition mechanism based on graph neural network is implemented. First, a dedicated R & D element ontology library is established, including hierarchical descriptions of key elements such as R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D results. The model matches the document content with the ontology library and performs message passing through the graph convolutional network to achieve accurate recognition of elements. The state update formula for each node is: hi^(t + 1) = σ(W·AGG({hj^t: j∈N(i)}), where hi represents the feature of node i, W is the weight matrix, AGG is the aggregation function, and N(i) is the neighbor set of node i.
[0040] In this embodiment, a cross-modal semantic association analysis method is innovatively designed. First, a multi-level semantic similarity calculation framework is constructed, including similarity metrics at the word level, sentence level, and paragraph level. At the word level, cosine similarity is used to calculate the similarity between feature vectors; at the sentence level, an improved mutual information calculation method is used to evaluate the degree of semantic association; at the paragraph level, a hierarchical attention network is used to capture long-distance dependencies.
[0041] In this embodiment, a dynamic graph update strategy is adopted when constructing the multimodal relationship graph. The initial graph is constructed based on the identified R & D elements. The nodes represent different element entities, and the edges represent the relationships between the elements. With the input of new information, the system dynamically updates the attributes of the nodes and edges through the graph attention network. The relationship strength calculation formula is: rij = MLP(concat[hi, hj, hi⊙hj]), where hi and hj are node features, ⊙ represents element-wise multiplication, and MLP is a multi-layer perceptron.
[0042] In this embodiment, a template-based document reorganization mechanism is designed. First, a standardized R & D document template library is established, which contains the structure templates of different types of R & D documents. The reorganization process adopts a graph-to-sequence generation method to reorganize the information in the multimodal relationship graph according to the template structure. The generation process uses an improved Transformer decoder, which adds structural constraint attention on the basis of the standard attention mechanism to ensure that the generated content conforms to the template specifications.
[0043] In this embodiment, an intelligent fusion strategy for multimodal information is implemented. During the document reorganization process, the system dynamically determines the display method of information according to the semantic importance of the content and the degree of support of multimodal evidence. For example, for key experimental data, the system will retain both the text description, supplementary explanations in the voice recording, and image evidence to form a complete information chain. Through the hierarchical attention mechanism, the system can accurately identify the importance weights of different modal information, ensuring that the fused document is both complete and concise.
[0044] In this embodiment, a document consistency check mechanism is also implemented. When generating the final document, the system performs consistency verification on cross-modal information, including numerical consistency, logical consistency, and temporal consistency. By establishing multi-dimensional check rules, the system can discover and mark potential information conflicts to ensure the accuracy and reliability of the generated document.
[0045] Through the above technological innovations, this embodiment effectively solves the problems in traditional R & D document processing, such as inaccurate element recognition, insufficient information association, and non-standard document reorganization. In practical applications, this solution can accurately identify the key elements in R & D documents, establish semantic associations between multimodal information, and generate standardized R & D documents. It is particularly suitable for document processing in complex R & D projects, significantly improving the intelligence level and processing efficiency of document processing. The adaptive characteristics of this solution enable it to process R & D documents in different fields, with broad application value.
[0046] As can be seen from the above description, the R & D document processing method provided by the embodiments of the present application can achieve unified processing of text, voice, and image information by constructing a multi-modal document parsing engine. Through cross-modal preprocessing and semantic alignment network, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify R & D elements, establish a multi-modal relationship graph, and realize intelligent extraction and correlation analysis of various types of information in R & D documents. This method effectively solves the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provides strong support for R & D process management and knowledge asset accumulation, and significantly improves the standardized management level of R & D documents.
[0047] In an embodiment of the R & D document processing method of the present application, it may specifically include the following content: Step S201: Construct a document format recognition network based on a neural network, extract a document feature vector from the document binary stream, use a convolutional neural network to classify and train the document feature vector to obtain a document format classification model, identify the format type of the input document based on the document format classification model, select a corresponding parsing rule library according to the recognition result to parse the document structure, and associate and save the parsed document structure information with the original document data; Step S202: Construct a voice feature extraction network and an image feature extraction network, use Mel-frequency cepstral coefficients to extract acoustic features from the R & D process recordings, input the acoustic features into a speech recognition model trained based on a long short-term memory network for text conversion, perform convolutional feature extraction and residual network training on experimental pictures and equipment photos to obtain a scene classification model, classify and label the picture content through the scene classification model, and integrate the converted voice text and picture annotation information with the document structure information to obtain multi-modal document data.
[0048] Optionally, in this embodiment, a format recognition network specifically for R & D documents is first constructed. Considering the diversity of R & D documents, a two-stream network structure is designed to process the visual features and content features of the documents respectively. The visual stream uses an improved VGG16 network to extract document layout features through five convolutional blocks. Each convolutional block contains two 3×3 convolutional layers and one max-pooling layer, and the number of channels in the convolutional layers gradually increases from 64 to 512. In order to improve the recognition ability of document format details, a spatial attention mechanism is introduced in the third and fourth convolutional blocks.
[0049] In this embodiment, an adaptive feature extraction method is innovatively designed in the document binary stream processing link. First, the document byte stream is scanned through a sliding window, and the window size is dynamically adjusted according to the document type. For the data within each window, the byte frequency distribution and n-gram features are calculated to construct an initial feature vector. Then, a multi-layer perceptron is used to perform dimensionality reduction and non-linear transformation on the features. The dimension of the features after dimensionality reduction is 256, which retains the key information of the document format.
[0050] In this embodiment, a format classification model based on a deep residual network is implemented. The backbone of the network adopts the ResNet50 structure, and the problem of difficult training of deep networks is effectively solved through residual connections. In the training stage, a multi-task learning strategy is adopted to predict the document format type and the document structure level simultaneously. The loss function is designed as: L = α·Lformat + β·Lstructure, where Lformat is the format classification loss, Lstructure is the structure prediction loss, and α and β are balance factors.
[0051] In this embodiment, an intelligent parsing rule selection mechanism is designed. A rule library containing common formats such as Word, PDF, Excel, etc. is established, and each format corresponds to specific structure parsing rules. The rule selection adopts a fuzzy inference system, comprehensively considering the confidence of format recognition and the document content features. For cases with low recognition confidence, the system will try parsing rules for multiple similar formats and determine the final parsing scheme by verifying the consistency of the results.
[0052] In this embodiment, in terms of speech feature extraction, an improved MFCC feature extraction method is adopted. First, the audio signal is pre-emphasized and framed, with a frame length of 25 ms and a frame shift of 10 ms. Then, 40-dimensional MFCC features are extracted through a Mel filter bank, and the first-order and second-order difference features are calculated to obtain a 120-dimensional acoustic feature vector. To improve the robustness of the features, a noise suppression algorithm based on spectral subtraction is introduced.
[0053] In this embodiment, a speech recognition model based on Transformer is innovatively implemented. The encoder adopts a multi-layer Transformer structure, and each layer contains a multi-head self-attention mechanism and a feed-forward neural network. The number of attention heads is set to 8, and the hidden layer dimension is 512. The decoder integrates a CTC-based acoustic model and a Transformer-based language model, and realizes the dynamic fusion of acoustic features and language features through the attention mechanism. The model is pre-trained on a professional field speech dataset, which improves the recognition accuracy of professional terms.
[0054] In this embodiment, in the image feature extraction stage, a multi-scale feature extraction network is designed. Based on ResNet101, the backbone network is constructed, and multi-scale feature fusion is achieved through the Feature Pyramid Network. At each scale, the attention mechanism is used to highlight the features of important regions. The network outputs a 2048-dimensional feature vector, which contains the local details and global semantic information of the image.
[0055] This embodiment implements an innovative scene classification model. The DenseNet161 architecture is adopted, and feature reuse is enhanced through dense connections, improving the feature utilization efficiency of the model. In the training stage, the knowledge distillation mechanism is introduced, and a pre-trained large model is used to guide the training of the scene classification model. The classification head adopts a multi-label classification strategy to predict scene categories and scene attributes simultaneously.
[0056] This embodiment designs a multi-modal information integration framework. The Graph Attention Network is used to establish the association between different modal information. The nodes represent different information units, and the edges represent the semantic relationships between the information. Through the message passing mechanism, cross-modal feature fusion is achieved. The fused features are selectively updated through the gated update unit, ensuring the accuracy of information integration.
[0057] This embodiment innovatively implements a document structure association preservation mechanism. A hierarchical index structure is established, which includes three levels: the format layer, the structure layer, and the content layer. The document type information is recorded in the format layer, the hierarchical structure of the document is saved in the structure layer, and the specific multi-modal information is stored in the content layer. Through the establishment of bidirectional links, fast retrieval of structural information and raw data is achieved.
[0058] Through the above technological innovations, this embodiment effectively solves problems such as inaccurate format recognition, incomplete feature extraction, and insufficient multi-modal information integration in the processing of R & D documents. In practical applications, this solution can accurately identify the format types of various R & D documents, complete speech transcription and image content extraction with high quality, and achieve effective integration of multi-modal information. It is particularly suitable for the document processing of complex R & D projects, significantly improving the automation level and accuracy of document processing. The modular design of this solution also facilitates optimization and expansion according to the characteristics of different R & D fields.
[0059] In an embodiment of the R & D document processing method of this application, the following content may also be specifically included: Step S301: Establish a document segmentation processing engine, segment the document content based on the document structure information, construct a text classification model using a bidirectional long short-term memory network, classify the themes of the segmented text fragments, identify the theme categories of each text fragment, segment the speech text according to the timestamp information, calculate the semantic similarity of the segmented speech text fragments, and merge the speech text fragments with a semantic similarity higher than the preset threshold to generate speech text paragraphs; Step S302: Construct a cross-modal feature fusion network to semantically match the text segment theme information with the speech text paragraph, extract keywords from the image annotation information, calculate the semantic correlation degree between the text segment, the speech paragraph and the image annotation based on the attention mechanism, and use the multi-head attention mechanism to perform weighted fusion on the semantic correlation degree to generate a fused multi-modal document feature matrix. Combine the multi-modal document feature matrix with the document structure information to form the original data containing multi-modal information.
[0060] Optionally, in this embodiment, an intelligent document segmentation processing engine is first designed according to the characteristics of R & D documents. Based on the title hierarchy, paragraph markers and format identifiers in the document structure information, a hierarchical segmentation strategy is adopted. During the segmentation process, a dynamic window mechanism based on semantic coherence is introduced, and the window size is adaptively adjusted according to the semantic integrity of the content. For paragraphs containing special elements such as formulas and tables, the system maintains their integrity through special markers to ensure that the segmentation does not damage the logical structure of the document.
[0061] This embodiment innovatively implements a text classification model based on BiLSTM. The model uses a word embedding layer to convert the input text into a dense vector representation with an embedding dimension of 300. The BiLSTM layer contains LSTM units in two directions, and the hidden layer dimension of each direction is 256. To improve the recognition ability of professional terms, a domain dictionary matching layer is added before the word embedding layer, and the feature representation of professional terms is strengthened through the attention mechanism. The classification task of the model adopts a multi-label strategy, and the output layer uses the sigmoid activation function, so that each text segment can belong to multiple relevant topic categories at the same time.
[0062] In the speech text processing link of this embodiment, an intelligent segmentation algorithm based on timestamps is designed. First, the initial segmentation points are determined according to the prosodic features such as pauses and intonation changes in the speech. Then, combined with the speaker recognition result, segmentation markers are added at the speaker switch. For segments that are too short, the system will perform merging processing based on the context semantics. The segmented speech text extracts semantic features through the BERT model, and calculates the cosine similarity between adjacent segments.
[0063] This embodiment implements an innovative method for generating speech text paragraphs. First, a semantic similarity calculation model is established, and a hierarchical similarity measurement strategy is adopted. At the word level, Word2Vec is used to calculate the word meaning similarity; at the sentence level, semantic vectors are extracted through a sentence encoder; at the paragraph level, an attention pooling mechanism is used to aggregate sentence features. The similarity threshold is dynamically adjusted through the validation set to ensure that the merged paragraphs maintain both semantic coherence and specific details.
[0064] This embodiment designs a multi-level cross-modal feature fusion network. The network adopts a three-stream architecture to process text, speech, and image features respectively. In the text stream, the BERT model is used to extract the semantic features of text segments; in the speech stream, the Transformer encoder is adopted to process the speech text features; in the image stream, the CLIP model is used to extract the visual-language joint representation of images. The features of the three streams perform feature interaction and alignment through the cross-attention mechanism.
[0065] This embodiment innovatively implements a semantic matching mechanism based on a graph neural network. The text segments and speech paragraphs are used as the nodes of the graph, and the edges are connected based on semantic similarity. Message passing is performed through the graph attention network to update the semantic representation of the nodes. The node feature update formula is: hi' = σ(Σj∈Ni αij·Whj), where hi' is the updated node feature, αij is the attention weight, W is the parameter matrix, and Ni is the set of neighbors of node i.
[0066] In the keyword extraction step of this embodiment, an improved TextRank algorithm is adopted. First, a word graph network is constructed, with the candidate keywords as the nodes, and the edge weights are calculated through the word co-occurrence relationship. Then, the importance scores of the nodes are calculated iteratively: S(Vi)= (1-d) + d·Σj∈In(Vi) wji·S(Vj) / Σj∈Out(Vj) wjk, where d is the damping coefficient, wji is the edge weight, and In(Vi) and Out(Vi) represent the sets of edges pointing to and starting from node Vi respectively.
[0067] This embodiment designs an innovative multi-head attention fusion mechanism. The number of attention heads is set to 8, and each attention head is responsible for capturing the correlation relationships at different semantic levels. For each attention head, the attention weight between the query vector Q, the key vector K, and the value vector V is calculated: Attention(Q,K,V) = softmax(QK^T / √dk)V, where dk is the dimension of the key vector. The outputs of multiple attention heads are fused through linear transformation and concatenation operations.
[0068] This embodiment implements a dynamic feature matrix generation method. The fused features are adjusted in dimension through an adaptive pooling layer to generate a feature matrix with a unified dimension. Each row of the feature matrix corresponds to a multi-modal information unit, which contains the unified representation of text semantics, speech features, and image semantics. To preserve the structural information of the document, position encoding and hierarchical markers are added to the feature matrix.
[0069] Through the above technological innovations, this embodiment effectively solves the problems of inaccurate content segmentation, insufficient multi-modal feature fusion, and inaccurate semantic association establishment in R & D document processing. In practical applications, this solution can accurately identify the semantic structure of documents, realize intelligent segmentation and fusion of multi-modal information, and significantly improve the intelligent level of document processing. It is particularly suitable for processing complex R & D project documents and can effectively maintain the semantic integrity and structure of documents. The adaptive characteristics of this solution enable it to process R & D documents in different fields and have broad application value.
[0070] In an embodiment of the R & D document processing method of this application, it may specifically include the following content: Step S401: Construct a block division model, establish information type recognition rules based on the structure marking information in the original data, use the conditional random field algorithm to perform sequence annotation on the original data, identify the boundary positions of text, speech, and images, divide the original data into independent blocks according to the boundary positions, establish a text word segmentation model based on word vectors, perform word segmentation processing on the text blocks, use the hidden Markov model for part-of-speech tagging, and input the tagging results into a pre-trained text encoder to generate text semantic vectors; Step S402: Construct a multi-modal feature extraction network, use a voiceprint recognition algorithm to extract the speaker's acoustic features for the voice block, train a speaker recognition model based on the acoustic features, segment the voice signal according to the silent segments, extract acoustic feature parameters for the segmented voice segments, convert the acoustic feature parameters into voice feature vectors through a recurrent neural network encoder, apply an object detection network to the image block to identify key objects in the image, use a deep residual network to classify the scene of the image, and fuse the object detection results and the scene classification results to generate image feature vectors.
[0071] Optionally, this embodiment first designs a special block division model for the complex structure of R & D documents. This model adopts a hierarchical processing architecture and first establishes a preliminary structure division based on the visible marks of the document (such as titles, paragraph marks, chart marks, etc.). On this basis, a rule-based information type recognizer is introduced to determine the type attributes of information by analyzing content features (such as text format, numerical distribution, multimedia marks, etc.). The rule library contains a special rule set for R & D documents, such as experimental data area recognition rules, chart description recognition rules, etc.
[0072] This embodiment innovatively implements a sequence labeling mechanism based on Conditional Random Field (CRF). The feature function of the CRF model consists of two parts: node features and edge features. The node features reflect the labeling probability of a single position, and the edge features describe the labeling transition relationship between adjacent positions. The design of the feature function takes into account multiple levels of information: f(yt, yt-1, xt) = w1·f1(yt, xt) + w2·f2(yt, yt-1), where yt represents the labeling at the current position, xt represents the observation feature, and w1 and w2 are weight parameters.
[0073] In the boundary recognition stage of this embodiment, a two-way scanning strategy is adopted. First, scan from front to back to identify possible starting positions of blocks; then scan from back to front to determine the ending positions of blocks. During the scanning process, the selection of boundary positions is optimized through a dynamic programming algorithm to ensure the rationality of the division result. The optimization objective consists of three components: intra-block consistency, inter-block difference, and structural integrity.
[0074] This embodiment implements an innovative text segmentation model. The model is based on an improved Word2Vec architecture, uses the Skip-gram algorithm to train word vectors, and the window size is dynamically adjusted. A larger window is used for professional terms to capture more context information. The dimension of the word vector is set to 300, and the training efficiency is optimized through negative sampling. The model is pre-trained on a professional domain corpus, significantly improving the recognition ability of professional terms.
[0075] In the part-of-speech tagging stage of this embodiment, an improved Hidden Markov Model (HMM) is adopted. The state space includes an extension of the standard part-of-speech set, adding special part-of-speech tags for R & D documents, such as technical indicator words, parameter nouns, etc. The transition probability matrix is obtained through training on a professional document corpus, and the emission probability takes into account multiple features such as word form and position. The optimal labeling sequence is solved through the Viterbi algorithm.
[0076] This embodiment designs a BERT-based text encoder. The encoder adopts a 12-layer Transformer structure, with a hidden layer dimension of 768 and 12 attention heads. In the pre-training stage, in addition to the regular masked language model task, a domain adaptation task is added to improve the model's understanding ability of professional texts. When extracting features, a multi-layer feature fusion strategy is adopted to integrate semantic information from different layers.
[0077] In terms of speech processing, this embodiment implements a high-precision speaker recognition algorithm. First, 60-dimensional MFCC features are extracted, including fundamental frequency features and their first and second order differences. Then, i-vector features are extracted through a deep network, which consists of 5 fully connected layers, each followed by batch normalization and the PReLU activation function. Finally, speaker similarity calculation is performed through PLDA.
[0078] In this embodiment, a voice segmentation algorithm based on energy and zero-crossing rate is innovatively designed. The voice activity region is identified through a dual-threshold detection mechanism, and the thresholds are adaptively adjusted according to the background noise level. To improve the accuracy of segmentation, a boundary optimization strategy based on prosodic features is introduced, considering factors such as pitch and energy changes.
[0079] In this embodiment, a voice feature encoder based on BiLSTM is implemented. The network consists of 3 layers of bidirectional LSTM with a hidden layer dimension of 256, and residual connections are used to improve gradient propagation. The encoder outputs a feature vector of a fixed dimension, retaining the acoustic features and semantic information of the voice. At the same time, the features of important time frames are highlighted through an attention mechanism.
[0080] In the image processing section of this embodiment, an improved YOLOv5 object detection network is adopted. The backbone feature extractor of the network uses the CSPDarknet53 structure, and spatial pyramid pooling is used to enhance the multi-scale object detection ability. A dedicated data augmentation strategy, such as random cropping and brightness adjustment, is designed according to the characteristics of the R & D scenario to improve the generalization ability of the model.
[0081] In this embodiment, an innovative feature fusion mechanism is designed. The features of object detection and scene classification are integrated through an adaptive feature fusion module, which includes two branches: channel attention and spatial attention. The calculation formula for the fused feature is: F = α·Fd + β·Fs, where Fd is the detection feature, Fs is the classification feature, and α and β are adaptive weight coefficients.
[0082] Through the above technological innovations, this embodiment effectively solves the problem of multi-modal data analysis in R & D documents. This solution can accurately divide different types of information blocks, precisely extract text, voice, and image features, and achieve effective fusion of features. In practical applications, it significantly improves the accuracy of subsequent content understanding and knowledge extraction. The modular design of the solution also supports optimization and expansion according to the characteristics of different R & D fields.
[0083] In an embodiment of the R & D document processing method of this application, the following content may also be specifically included: Step S501: Construct a cross-modal feature mapping network, train a feature conversion model using an adversarial learning method, input the text semantic vector, voice feature vector, and image feature vector into the corresponding feature conversion models respectively to generate feature representations with the same dimension, calculate the mutual information loss for the converted feature representations, and optimize the parameters of the feature conversion model through backpropagation so that different modal features have similar distribution characteristics in the unified feature space, and calculate the semantic similarity matrix between different modal features based on cosine similarity; Step S502: Construct a feature alignment optimization network, establish feature alignment constraints based on the semantic similarity matrix, perform relational reasoning on the feature representations using a graph neural network, perform attention-weighted fusion on the reasoning results and the feature representations, perform dimensionality reduction on the fused features through a multi-layer perceptron, normalize the features after dimensionality reduction to obtain a standardized representation, establish a mapping relationship between the standardized representations with corresponding relationships and the original features, and generate standardized document data containing multi-modal alignment information.
[0084] Optionally, in this embodiment, an innovative cross-modal feature mapping network is first constructed according to the characteristics of R & D documents. This network adopts a three-branch structure, and each branch contains a feature transformer and a discriminator. The feature transformer is designed based on the Transformer architecture and contains 6 encoder layers, and each layer contains a multi-head self-attention mechanism and a feed-forward neural network. In order to handle the feature differences of different modalities, a modality-specific embedding layer is designed at the input layer to map features of different dimensions to a unified 512-dimensional feature space.
[0085] This embodiment innovatively implements an adversarial learning training strategy. The discriminator network adopts a multi-layer convolutional structure and attempts to distinguish the feature distributions after different modality conversions. The feature transformer, through adversarial training, generates feature representations that can deceive the discriminator. The adversarial loss function is designed as: Ladv = E[log D(x)] + E[log(1 - D(G(z)))]), where D is the discriminator, G is the feature transformer, x is the sample of the true feature distribution, and z is the input feature. This adversarial mechanism promotes the feature distributions of different modalities to tend to be consistent in a unified space.
[0086] This embodiment designs a feature consistency constraint based on mutual information. A neural network-based mutual information estimator is used to calculate the mutual information between the features before and after conversion. The mutual information loss function is: LMI = -I(X;Y) = -E[log(P(X|Y) / P(X))], where X is the original feature and Y is the converted feature. By minimizing the mutual information loss, it is ensured that the semantic content of the original information is retained during the feature conversion process.
[0087] In this embodiment, an innovative cycle consistency checking mechanism is implemented during the feature conversion process. The converted features are reversely converted, and the reconstruction error with the original features is calculated. The cycle consistency loss is: Lcyc = ||G2(G1(x1)) - x1||2 + ||G1(G2(x2)) - x2||2, where G1 and G2 are the converters between different modalities. This mechanism ensures the reversibility and stability of the feature conversion.
[0088] In this embodiment, a feature alignment optimization network is innovatively designed. First, a feature relationship graph is constructed based on the semantic similarity matrix. The nodes in the graph represent feature representations of different modalities, and the edge weights are determined by the cosine similarity. The graph attention network is used for feature optimization, and the attention weight calculation formula is: αij = softmax(LeakyReLU(Wa[Whi||Whj])), where Wa and W are learnable parameter matrices, and hi and hj are node features.
[0089] This embodiment implements a multi-level relationship reasoning mechanism. A three-layer graph convolutional network is constructed, and the message passing function for each layer is: h'i = σ(Σj∈Ni αij·Whj / |Ni|), where Ni is the neighbor set of node i, and |Ni| is the number of neighbors. Through multi-layer transmission, features can fully fuse context information to form a richer semantic representation.
[0090] This embodiment designs an attention-weighted feature fusion strategy. Calculate the attention weight for the inference result of the graph neural network and the original feature representation: Attention(Q,K,V) = softmax(QK^T / √d)V, where Q is the query matrix, K is the key matrix, V is the value matrix, and d is the feature dimension. In this way, the system can adaptively select important features for fusion.
[0091] In the dimensionality reduction processing link of this embodiment, an improved multi-layer perceptron structure is adopted. The network contains three hidden layers with dimensions of 256, 128, and 64 in sequence. Each layer is followed by a batch normalization and a dropout layer. The activation function uses GELU, which has better non-linear characteristics than the traditional ReLU. Through this structure, both the discriminability of the features is maintained, and the computational complexity is significantly reduced.
[0092] This embodiment implements an innovative feature standardization method. Perform L2 normalization on the dimensionality-reduced features: x' = x / ||x||2 to ensure that the norms of different feature vectors are consistent. Then adjust the feature distribution through batch normalization: y = γ(x-μ) / √(σ2+ε) + β, where γ and β are learnable parameters, and μ and σ are batch statistics. This standardization process improves the stability and comparability of the features.
[0093] This embodiment designs a mechanism for establishing a bidirectional mapping relationship. A mapping table is maintained for each standardized feature to record its corresponding relationship with the original feature. The mapping relationship includes the feature index, transformation parameters, and confidence score. The system maintains the timeliness and accuracy of the mapping relationship through a regular update mechanism.
[0094] Through the above technological innovations in this embodiment, problems such as inaccurate cross-modal feature alignment, information loss in feature transformation, and incomplete establishment of semantic associations in R & D document processing are effectively solved. In practical applications, this solution can accurately achieve the unified representation and alignment of different modal features, maintain the semantic information of the original features, and significantly improve the accuracy and efficiency of multi-modal document processing. It is particularly suitable for document processing in complex R & D projects and can effectively handle the fusion and alignment of multi-modal data such as text, speech, and images. The adaptive characteristics of this solution enable it to process R & D documents in different fields and have broad application prospects.
[0095] In an embodiment of the R & D document processing method of this application, the following content may also be specifically included: Step S601: Construct a multi-modal feature sequence encoding network, use the position encoding method to encode the position information of the feature sequence in the standardized document data, input the encoded feature sequence into the encoder based on the Transformer architecture, perform multi-head self-attention calculation on the feature sequence to obtain the context representation, use the recurrent neural network to perform temporal modeling on the context representation, input the modeling result into the conditional random field for sequence annotation, and classify and annotate the document content based on the preset R & D element annotation system to identify the text segments corresponding to the R & D objectives, technical solutions, innovation points, R & D processes, experimental data, and R & D results; Step S602: Construct an element recognition optimization network, train an element classifier based on the pre-annotated R & D document dataset, input the sequence annotation result into the element classifier for multi-label classification, calculate the cross-entropy loss between the classification result and the preset label, use the gradient descent method to optimize the classifier parameters, re-extract features and perform classification prediction on the text segments with classification confidence lower than the threshold, and integrate the optimized classification result with the sequence annotation result to generate the final R & D element recognition result.
[0096] Optionally, in this embodiment, a dedicated multi-modal feature sequence encoding network is first designed. Aiming at the heterogeneity of text, speech, and image features in R & D documents, a unified position encoding method is used for feature alignment. The position encoding uses an improved sine position encoding function: PE(pos,2i) = sin(pos / 10000^(2i / d)), PE(pos,2i + 1) = cos(pos / 10000^(2i / d)), where pos represents the sequence position, i represents the dimension index, and d is the model dimension. This encoding method not only maintains the relative position information of the features but also reflects the hierarchical relationship of different modal information.
[0097] In this embodiment, a multi-layer Transformer encoder is innovatively implemented in the feature sequence processing stage. The encoder consists of 6 encoding layers, each of which is composed of a multi-head self-attention module and a feed-forward neural network. The number of attention heads is set to 8, enabling the model to capture the dependencies between features from different perspectives. The self-attention calculation adopts the scaled dot-product attention mechanism: Attention(Q, K, V) = softmax(QK^T / √dk)V, where Q, K, and V are the query, key, and value matrices respectively, and dk is the dimension of the key vector.
[0098] This embodiment designs an innovative temporal modeling mechanism. A bidirectional GRU network is used to process the context representation output by the Transformer, with a hidden layer dimension of 512. The gating mechanism of the GRU can effectively process long sequence information, and the design of the update gate and reset gate enables the model to dynamically adjust the information flow according to the context. At the same time, residual connections and layer normalization are introduced to improve the training effect of the deep network.
[0099] This embodiment implements a sequence annotation framework based on CRF. The CRF layer takes into account the transition constraints between labels, especially the logical relationships between R & D elements. For example, technical solutions usually follow R & D goals, and experimental data is often associated with the R & D process. The transition probability matrix is obtained through statistical learning, reflecting the typical organizational patterns of elements in R & D documents.
[0100] This embodiment innovatively constructs an R & D element annotation system. This system includes multiple levels: the first level is the basic element categories, such as R & D goals, technical solutions, etc.; the second level is the specific attributes of the elements, such as the innovativeness of the goal, the feasibility of the solution, etc.; the third level is the association relationships between elements, such as the goal-solution correspondence relationship, the process-result causal relationship, etc. Through this hierarchical annotation system, the content structure of R & D documents can be comprehensively grasped.
[0101] This embodiment designs an efficient element classifier. The classifier adopts a multi-task learning framework to simultaneously perform element category recognition and attribute prediction. The backbone network uses a pre-trained BERT model, which is fine-tuned with domain-specific corpus. The classification head adopts a multi-layer perceptron, and each task corresponds to an independent classification layer. The loss function is designed as a weighted sum of multiple task losses: L = Σi αi·Li, where αi is the task weight and Li is the cross-entropy loss of each task.
[0102] This embodiment implements an optimization mechanism based on confidence. For text segments with a classification confidence lower than the threshold, the system will initiate a secondary processing process. First, context features are re-extracted through the attention mechanism, and then an ensemble learning method is used for classification. The ensemble classifier includes random forest, XGBoost, and a deep neural network, and the final classification result is determined through a voting mechanism.
[0103] In this embodiment, an integration strategy for classification results is innovatively designed. The weighted average method is used to fuse the sequence annotation and classification prediction results, and the weight coefficients are dynamically adjusted through the validation set. For conflicting prediction results, a rule-based arbitration mechanism is introduced to ensure the consistency of the final results. At the same time, post-processing is performed through the knowledge graph to correct the prediction results that do not conform to domain knowledge.
[0104] In the feature extraction stage of this embodiment, an adaptive feature enhancement mechanism is implemented. According to the length, complexity, and domain characteristics of text segments, the feature extraction strategy is dynamically adjusted. For segments containing technical terms, the domain dictionary matching feature is added; for content related to experimental data, the extraction of numerical features is strengthened; for descriptions of innovation points, comparative and differential expressions are focused on.
[0105] Through these technological innovations in this embodiment, problems such as inaccurate identification, incomplete classification, and insufficient association of elements in R & D documents are effectively solved. This solution can accurately identify various R & D elements, establish semantic associations between elements, and ensure the reliability of the identification results. In practical applications, it significantly improves the intelligent processing level of R & D documents, providing a reliable basis for subsequent knowledge management and decision support. It is particularly suitable for document analysis in large R & D projects and can effectively extract and organize key R & D information. The adaptive characteristics of this solution enable it to process R & D documents in different fields and have broad application value.
[0106] In an embodiment of the R & D document processing method of this application, the following content may also be specifically included: Step S701: Construct a multi-modal relationship reasoning network, calculate the semantic similarity of the identified R & D elements to establish an association matrix between elements, perform sparsification processing on the association matrix based on the graph attention network to obtain an initial graph structure, use the feature representations of different modalities as the attribute information of graph nodes, perform feature aggregation on the graph structure using the message passing mechanism, update the semantic representations of nodes through a multi-layer graph convolutional network, calculate the relationship strength between nodes based on the updated semantic representations, and construct a multi-modal relationship graph containing text descriptions, voice records, and image evidence; Step S702: Construct a document restructuring and generation network, establish document structure constraints based on a preset document template, map the nodes in the multi-modal relationship graph according to the template structure, use a graph-to-sequence generation model to convert the graph structure into a text sequence, perform coherence optimization at the discourse level on the generated text sequence, insert the corresponding voice records and image evidence into the corresponding positions according to the semantic association relationship, and use a text polishing model to optimize the language of the inserted document content to generate R & D document data integrating multi-modal information.
[0107] Optionally, this embodiment first constructs a special multimodal relational reasoning network. For the R&D elements in the R&D documents, a semantic similarity calculation module based on BERT is designed. This module adopts a twin network structure, inputs the element pairs into the BERT encoder with shared weights, and extracts context-related semantic features through the attention mechanism. The similarity calculation adopts an improved cosine metric: sim(x,y) = cos(Wx,Wy), where W is a learnable projection matrix, and x and y are the semantic representations of the elements.
[0108] This embodiment innovatively implements the sparse processing of the association matrix. An adaptive threshold strategy is adopted, and the threshold value is dynamically adjusted according to the similarity distribution: threshold = μ + α·σ, where μ is the mean similarity, σ is the standard deviation, and α is an adjustable parameter. For associations below the threshold, the system will further verify through the knowledge graph, retain weak similarity connections with logical associations, and thus build a more accurate initial graph structure.
[0109] This embodiment designs a multimodal node attribute representation method. For text nodes, position encoding and semantic encoding are combined; for speech nodes, acoustic features and text features are fused; for image nodes, visual features and scene semantics are integrated. The attribute vector of the node is generated by a multimodal feature fusion network, and the weights of different modal features are dynamically adjusted using a gating mechanism: h = σ(Wg·[ht;hv;ha])·tanh(Wf·[ht;hv;ha]), where ht, hv, and ha are text, visual, and speech features, respectively.
[0110] This embodiment implements an innovative message transmission mechanism. A bidirectional message flow is designed on the graph structure. The forward flow transmits the dependency relationship between elements, and the reverse flow transmits the evidence support relationship. The message update function is: mi→j = fmsg(hi,hj,eij), where hi and hj are node features, eij is edge feature, and fmsg is a multi-layer perceptron. In order to improve the efficiency of message transmission, an attention gating mechanism is introduced to selectively receive important information.
[0111] In this embodiment, a multi-scale feature extraction strategy is adopted in the graph convolution network. The network contains three graph convolution layers, and the receptive field of each layer is gradually expanded to capture the semantic relationship from local to global. The node feature update formula is: h'i = σ(Σj∈Ni αij·(Whj + b) / |Ni|), where αij is the attention weight, W and b are learnable parameters, and Ni is the neighbor set of node i.
[0112] In this embodiment, a document reorganization and generation network is innovatively designed. Based on the specification requirements of R & D documents, a hierarchical document template library is established, which includes standard chapter structures such as introduction, method, experiment, and conclusion. The template constraints are implemented in a soft constraint manner, allowing for flexible adjustment according to the actual content. The constraint strength is determined by a confidence score: score = λ·template_match + (1-λ)·content_coherence.
[0113] In this embodiment, an efficient graph-to-sequence generation model is implemented. An encoder-decoder architecture based on Transformer is adopted. The encoder processes the graph structure information, and the decoder generates the target sequence. In order to maintain the integrity of multimodal information, a multimodal attention mechanism is introduced during the decoding process, which pays attention to text content, voice records, and image evidence simultaneously. The beam search strategy is used in the generation process to keep multiple candidate sequences.
[0114] In this embodiment, a coherence optimization method at the discourse level is designed. A discourse relationship network based on semantic role labeling is constructed. By identifying argument relationships, temporal relationships, and logical relationships, the connections between paragraphs are optimized. For the detected incoherences, the system will automatically generate transitional sentences or add conjunctions to ensure the fluency of the document.
[0115] In this embodiment, the intelligent insertion of multimodal information is innovatively realized. The system analyzes the semantic association strength between voice records and image evidence and the text content, and inserts multimodal information at appropriate positions. The selection of the insertion position takes into account context coherence and information integrity, and the optimal insertion scheme is solved through a dynamic programming algorithm. For the inserted content, the system will generate necessary explanatory text to ensure the readability of the document.
[0116] In this embodiment, a text polishing module based on a pre-trained language model is designed. This module is fine-tuned on professional field documents, can identify and correct non-standard expressions, optimize sentence structures, and ensure the accurate use of professional terms. The polishing process adopts an iterative optimization strategy, and different levels of optimization objectives are focused on in each iteration, including grammar accuracy, expression standardization, and professionalism.
[0117] Through the above technological innovations in this embodiment, problems such as unreasonable organization of multimodal information, non-standard document structure, and unprofessional expression in the generation of R & D documents are effectively solved. In practical applications, this solution can accurately establish the association relationships between R & D elements, reasonably organize multimodal information, and generate R & D documents that meet the specifications. It is particularly suitable for the automated generation of documents for complex R & D projects, significantly improving the quality and generation efficiency of documents. The adaptive characteristics of this solution enable it to handle R & D documents in different fields and have broad application value.
[0118] In order to effectively address the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provide strong support for R & D process management and knowledge asset accumulation, and significantly improve the standardized management level of R & D documents, this application provides an embodiment of an R & D document processing device for implementing all or part of the content of the R & D document processing method. See Figure 2 , the R & D document processing device specifically includes the following content: A multi-modal parsing module for establishing a multi-modal document parsing engine. The multi-modal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. The R & D document is input into the document format recognition module for format type recognition, the R & D process recording is input into the speech recognition module for speech-to-text processing, the experimental pictures and equipment photos are input into the image recognition module for feature extraction and scene recognition, and the content extraction module is used to extract text content, speech text, and image annotations to generate raw data containing multi-modal information; A cross-modal processing module for performing cross-modal preprocessing on the raw data. The raw data is divided into text blocks, speech blocks, and image blocks according to information types. The text blocks are segmented and part-of-speech tagged to construct text semantic vectors, the speech blocks are speaker-identified and speech-segmented to extract speech feature vectors, the image blocks are object-detected and scene-classified to generate image feature vectors, and the text semantic vectors, speech feature vectors, and image feature vectors are input into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space; A document standardization module for inputting the standardized document data into a multi-modal content understanding model trained based on deep learning. The R & D elements in the document are identified through the multi-modal content understanding model, including R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D results. Cross-modal semantic association analysis is performed on the identified R & D elements to establish a multi-modal relationship graph containing text descriptions, speech records, and image evidence. Based on the multi-modal relationship graph, the standardized document data is reorganized according to a unified document template format to generate R & D document data integrating multi-modal information.
[0119] As can be seen from the above description, the R & D document processing device provided by the embodiments of the present application can unify the processing of text, speech, and image information by constructing a multimodal document parsing engine. Through cross-modal preprocessing and semantic alignment network, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify R & D elements, establish a multimodal relationship graph, and realize the intelligent extraction and correlation analysis of various types of information in R & D documents. This method effectively solves the deficiencies of traditional technologies in multimodal information fusion and knowledge system construction, provides strong support for R & D process management and knowledge asset accumulation, and significantly improves the standardized management level of R & D documents.
[0120] At the hardware level, in order to effectively solve the deficiencies of traditional technologies in multimodal information fusion and knowledge system construction, provide strong support for R & D process management and knowledge asset accumulation, and significantly improve the standardized management level of R & D documents, the present application provides an embodiment of an electronic device for implementing all or part of the content in the R & D document processing method. The electronic device specifically includes the following: A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between the R & D document processing device and related devices such as a core business system, a user terminal, and a related database, etc. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the R & D document processing method and the embodiments of the R & D document processing device in the embodiments, and the content is incorporated herein, and the repeated parts will not be elaborated.
[0121] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0122] In practical applications, part of the R & D document processing method can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. The present application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.
[0123] The above-mentioned client device may have a communication module (i.e., a communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side. In other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0124] Figure 3 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 3 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to achieve telecommunication functions or other functions.
[0125] In one embodiment, the function of the R & D document processing method may be integrated into the central processing unit 9100. Among them, the central processing unit 9100 may be configured to perform the following controls: Step S101: Establish a multimodal document parsing engine. The multimodal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. Input the R & D document into the document format recognition module for format type recognition, input the R & D process recording into the speech recognition module for speech-to-text processing, input the experimental pictures and device photos into the image recognition module for feature extraction and scene recognition, and use the content extraction module to extract the text content, speech text, and image annotations to generate raw data containing multimodal information; Step S102: Perform cross-modal preprocessing on the raw data. Divide the raw data into text blocks, speech blocks, and image blocks according to the information type. Perform word segmentation and part-of-speech tagging on the text blocks to construct text semantic vectors, perform speaker recognition and speech segmentation on the speech blocks to extract speech feature vectors, perform object detection and scene classification on the image blocks to generate image feature vectors, and input the text semantic vectors, speech feature vectors, and image feature vectors into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space; Step S103: Input the standardized document data into a multi-modal content understanding model trained based on deep learning. Identify the R & D elements in the document through the multi-modal content understanding model, including R & D goals, technical solutions, innovation points, R & D processes, experimental data, and R & D achievements. Conduct cross-modal semantic association analysis on the identified R & D elements, establish a multi-modal relationship graph containing text descriptions, voice recordings, and image evidence, and reorganize the standardized document data according to a unified document template format based on the multi-modal relationship graph to generate R & D document data integrating multi-modal information.
[0126] As can be seen from the above description, the electronic device provided in the embodiment of the present application realizes the unified processing of text, voice, and image information by constructing a multi-modal document parsing engine. Through cross-modal preprocessing and semantic alignment network, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify R & D elements and establish a multi-modal relationship graph, realizing the intelligent extraction and correlation analysis of various types of information in R & D documents. This method effectively solves the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provides strong support for R & D process management and knowledge asset accumulation, and significantly improves the standardized management level of R & D documents.
[0127] In another implementation manner, the R & D document processing device can be separately configured from the central processing unit 9100. For example, the R & D document processing device can be configured as a chip connected to the central processing unit 9100, and the functions of the R & D document processing method are realized through the control of the central processing unit.
[0128] As Figure 3 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include all the components shown in Figure 3 ; in addition, the electronic device 9600 may further include components not shown in Figure 3 , and reference can be made to the prior art.
[0129] As Figure 3 shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.
[0130] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above-mentioned failure-related information can be stored, and in addition, a program for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to achieve information storage or processing, etc.
[0131] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.
[0132] The memory 9140 can be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be a memory that stores information even when power is off, can be selectively erased, and has more data. An example of this memory is sometimes referred to as an EPROM, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage unit 9142, which is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.
[0133] The memory 9140 can also include a data storage unit 9143, which is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 can include various drivers for the communication function of the electronic device and / or for executing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0134] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.
[0135] Based on different communication technologies, in the same electronic device, multiple communication modules 9110 can be provided, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 can include any suitable buffers, decoders, amplifiers, etc. Additionally, the audio processor 9130 is also coupled to a central processor 9100, so that it is possible to record on the device through the microphone 9132 and play the sound stored on the device through the speaker 9131.
[0136] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the research and development document processing method with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the research and development document processing method with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented: Step S101: Establish a multimodal document parsing engine. The multimodal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module, and a content extraction module. Input the research and development document into the document format recognition module for format type recognition, input the recording during the research and development process into the speech recognition module for speech-to-text processing, input the experimental pictures and device photos into the image recognition module for feature extraction and scene recognition, and use the content extraction module to extract the text content, speech text, and image annotations to generate raw data containing multimodal information. Step S102: Perform cross-modal preprocessing on the raw data. Divide the raw data into text blocks, speech blocks, and image blocks according to the information type. Perform word segmentation and part-of-speech tagging on the text blocks to construct text semantic vectors, perform speaker recognition and speech segmentation on the speech blocks to extract speech feature vectors, perform object detection and scene classification on the image blocks to generate image feature vectors, and input the text semantic vectors, speech feature vectors, and image feature vectors into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space. Step S103: Input the standardized document data into a multi-modal content understanding model trained based on deep learning. Identify the R & D elements in the document through the multi-modal content understanding model, including R & D objectives, technical solutions, innovation points, R & D processes, experimental data, and R & D achievements. Conduct cross-modal semantic association analysis on the identified R & D elements, establish a multi-modal relationship graph containing text descriptions, voice recordings, and image evidence, and reorganize the standardized document data according to a unified document template format based on the multi-modal relationship graph to generate R & D document data integrating multi-modal information.
[0137] As can be seen from the above description, the computer-readable storage medium provided by the embodiments of the present application realizes the unified processing of text, voice, and image information by constructing a multi-modal document parsing engine. Through cross-modal preprocessing and a semantic alignment network, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify R & D elements and establish a multi-modal relationship graph, realizing the intelligent extraction and association analysis of various types of information in R & D documents. This method effectively solves the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provides strong support for R & D process management and knowledge asset accumulation, and significantly improves the standardized management level of R & D documents.
[0138] The embodiments of the present application also provide a computer program product capable of implementing all the steps in the R & D document processing method with the execution subject being a server or a client in the above embodiments. When the computer program / instructions are executed by a processor, the steps of the R & D document processing method are implemented. For example, the computer program / instructions implement the following steps: Step S101: Establish a multi-modal document parsing engine. The multi-modal document parsing engine includes a document format recognition module, a voice recognition module, an image recognition module, and a content extraction module. Input the R & D document into the document format recognition module for format type recognition, input the R & D process recording into the voice recognition module for speech-to-text processing, input the experimental pictures and equipment photos into the image recognition module for feature extraction and scene recognition, and use the content extraction module to extract text content, voice text, and image annotations to generate raw data containing multi-modal information. Step S102: Conduct cross-modal preprocessing on the raw data. Divide the raw data into text blocks, voice blocks, and image blocks according to the information type. Perform word segmentation and part-of-speech tagging on the text blocks to construct text semantic vectors, perform speaker recognition and speech segmentation on the voice blocks to extract voice feature vectors, perform object detection and scene classification on the image blocks to generate image feature vectors, and input the text semantic vectors, voice feature vectors, and image feature vectors into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space. Step S103: Input the standardized document data into a multi-modal content understanding model trained based on deep learning. Identify the R & D elements in the document through the multi-modal content understanding model, including R & D objectives, technical solutions, innovation points, R & D processes, experimental data, and R & D achievements. Conduct cross-modal semantic association analysis on the identified R & D elements, establish a multi-modal relationship graph containing text descriptions, voice recordings, and image evidence, and reorganize the standardized document data according to a unified document template format based on the multi-modal relationship graph to generate R & D document data integrating multi-modal information.
[0139] As can be seen from the above description, the computer program product provided by the embodiments of the present application realizes the unified processing of text, voice, and image information by constructing a multi-modal document parsing engine. Through cross-modal preprocessing and semantic alignment networks, different types of data are converted into standardized feature representations. The system uses a deep learning model to identify R & D elements and establish a multi-modal relationship graph, realizing the intelligent extraction and association analysis of various types of information in R & D documents. This method effectively solves the deficiencies of traditional technologies in multi-modal information fusion and knowledge system construction, provides strong support for R & D process management and knowledge asset accumulation, and significantly improves the standardized management level of R & D documents.
[0140] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, apparatus, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0144] In the present invention, specific embodiments are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A research and development document processing method, characterized in that: The method comprises: Establish a multimodal document parsing engine, which includes a document format recognition module, a voice recognition module, an image recognition module and a content extraction module. The R&D document is input into the document format recognition module for format type recognition, the recording of the R&D process is input into the voice recognition module for voice-to-text processing, the experimental pictures and equipment photos are input into the image recognition module for feature extraction and scene recognition, and the content extraction module is used to extract text content, voice text, and image annotations to generate raw data containing multimodal information; Performing cross-modal preprocessing on the raw data, dividing the raw data into text blocks, voice blocks and image blocks according to information type, performing word segmentation and part-of-speech tagging on the text blocks, constructing text semantic vectors, performing speaker recognition and voice segmentation on the voice blocks, extracting voice feature vectors, performing target detection and scene classification on the image blocks, generating image feature vectors, inputting the text semantic vectors, voice feature vectors and image feature vectors into a cross-modal alignment network for semantic matching, and generating standardized document data in a unified feature space; The standardized document data is input into a multimodal content understanding model based on deep learning training, and the R&D elements in the document are identified through the multimodal content understanding model, including R&D goals, technical solutions, innovations, R&D processes, experimental data and R&D results. Cross-modal semantic association analysis is performed on the identified R&D elements, and a multimodal relationship map containing text descriptions, voice records and image evidence is established. Based on the multimodal relationship map, the standardized document data is reorganized in a unified document template format to generate R&D document data that integrates multimodal information.
2. The R&D document processing method according to claim 1, characterized in that: The multimodal document parsing engine is established, and the multimodal document parsing engine includes a document format recognition module, a speech recognition module, an image recognition module and a content extraction module. The R&D document is input into the document format recognition module for format type recognition, the R&D process recording is input into the speech recognition module for speech-to-text processing, and the experimental pictures and equipment photos are input into the image recognition module for feature extraction and scene recognition, including: Construct a document format recognition network based on a neural network, extract document feature vectors from document binary streams, use a convolutional neural network to classify and train the document feature vectors to obtain a document format classification model, identify the format type of the input document based on the document format classification model, select a corresponding parsing rule library to parse the document structure according to the recognition result, and associate the parsed document structure information with the original document data for storage; A speech feature extraction network and an image feature extraction network are constructed, and the Mel-frequency cepstral coefficients are used to extract acoustic features from recordings of the R&D process. The acoustic features are input into a speech recognition model trained based on a long short-term memory network for text conversion. Convolutional feature extraction and residual network training are performed on experimental images and equipment photos to obtain a scene classification model. The image content is classified and annotated through the scene classification model, and the converted speech text and image annotation information are integrated with the document structure information to obtain multimodal document data.
3. The R&D document processing method according to claim 1, characterized in that: The content extraction module is used to extract text content, voice text, and image annotations to generate raw data containing multimodal information, including: Establish a document segmentation processing engine, segment the document content based on the document structure information, use a bidirectional long short-term memory network to build a text classification model, classify the segmented text segments by topic, identify the topic category of each text segment, segment the speech text according to the timestamp information, calculate the semantic similarity of the segmented speech text, merge the speech text segments with semantic similarity higher than the preset threshold, and generate speech text paragraphs; A cross-modal feature fusion network is constructed to semantically match the topic information of text fragments with the voice text paragraphs, extract keywords from the image annotation information, calculate the semantic correlation between text fragments, voice paragraphs and image annotations based on the attention mechanism, and use a multi-head attention mechanism to weightedly fuse the semantic correlation to generate a fused multimodal document feature matrix. The multimodal document feature matrix is combined with the document structure information to form original data containing multimodal information.
4. The R&D document processing method according to claim 1, characterized in that: The method of dividing the original data into text blocks, voice blocks and image blocks according to information type, performing word segmentation and part-of-speech tagging on the text blocks, constructing text semantic vectors, performing speaker recognition and voice segmentation on the voice blocks, extracting voice feature vectors, performing target detection and scene classification on the image blocks, and generating image feature vectors includes: Construct a block partitioning model, establish information type recognition rules based on the structural tag information in the original data, use the conditional random field algorithm to sequence the original data, identify the boundary positions of text, speech and images, divide the original data into independent blocks according to the boundary positions, establish a text segmentation model based on word vectors, perform word segmentation on text blocks, use the hidden Markov model to perform part-of-speech tagging, and input the tagging results into the pre-trained text encoder to generate text semantic vectors; A multimodal feature extraction network is constructed, and a voiceprint recognition algorithm is used to extract the speaker's acoustic features for the speech block. The speaker recognition model is trained based on the acoustic features. The speech signal is segmented according to silent segments, and acoustic feature parameters are extracted from the segmented speech segments. The acoustic feature parameters are converted into speech feature vectors through a recursive neural network encoder. The target detection network is applied to the image block to identify the key targets in the image. The deep residual network is used to perform scene classification on the image. The target detection results are fused with the scene classification results to generate an image feature vector.
5. The R&D document processing method according to claim 1, characterized in that: The step of inputting the text semantic vector, the speech feature vector and the image feature vector into a cross-modal alignment network for semantic matching to generate standardized document data in a unified feature space includes: Construct a cross-modal feature mapping network, use adversarial learning to train the feature conversion model, input the text semantic vector, speech feature vector and image feature vector into the corresponding feature conversion model respectively, generate feature representations with the same dimension, calculate the mutual information loss for the converted feature representations, optimize the parameters of the feature conversion model through back propagation, make the different modal features have similar distribution characteristics in the unified feature space, and calculate the semantic similarity matrix between the different modal features based on cosine similarity; A feature alignment optimization network is constructed, feature alignment constraints are established based on the semantic similarity matrix, graph neural networks are used to perform relational reasoning on feature representations, the reasoning results are fused with feature representations in an attention-weighted manner, the fused features are reduced in dimension through a multi-layer perceptron, the reduced features are normalized to obtain a standardized representation, a mapping relationship is established between the standardized representations with corresponding relationships and the original features, and standardized document data containing multimodal alignment information is generated.
6. The R&D document processing method according to claim 1, characterized in that: The step of inputting the standardized document data into a multimodal content understanding model based on deep learning training, and identifying R&D elements in the document through the multimodal content understanding model, including R&D goals, technical solutions, innovations, R&D processes, experimental data and R&D results, includes: Construct a multimodal feature sequence encoding network, use the position encoding method to encode the position information of the feature sequence in the standardized document data, input the encoded feature sequence into the encoder based on the Transformer architecture, perform multi-head self-attention calculation on the feature sequence to obtain the context representation, use a recurrent neural network to perform time series modeling on the context representation, input the modeling results into the conditional random field for sequence labeling, classify and label the document content based on the preset R&D element labeling system, and identify the text fragments corresponding to the R&D goals, technical solutions, innovations, R&D processes, experimental data, and R&D results; Construct a feature recognition optimization network, train a feature classifier based on a pre-labeled R&D document dataset, input the sequence labeling results into the feature classifier for multi-label classification, calculate the cross entropy loss between the classification results and the preset labels, use the gradient descent method to optimize the classifier parameters, re-extract features and classify text fragments whose classification confidence is lower than the threshold, integrate the optimized classification results with the sequence labeling results, and generate the final R&D feature recognition results.
7. The R&D document processing method according to claim 1, characterized in that: The cross-modal semantic association analysis is performed on the identified R&D elements to establish a multimodal relationship map including text descriptions, voice records and image evidence, and the standardized document data is reorganized according to a unified document template format based on the multimodal relationship map to generate R&D document data integrating multimodal information, including: Construct a multimodal relational reasoning network, use semantic similarity calculation to establish the association matrix between the identified R&D elements, perform sparse processing on the association matrix based on the graph attention network to obtain the initial graph structure, use the feature representations of different modalities as the attribute information of the graph nodes, use the message passing mechanism to aggregate features on the graph structure, update the semantic representation of the nodes through a multi-layer graph convolutional network, calculate the relationship strength between the nodes based on the updated semantic representation, and construct a multimodal relational graph containing text descriptions, voice records, and image evidence; Construct a document reorganization generation network, establish document structure constraints based on preset document templates, map the nodes in the multimodal relationship graph according to the template structure, use the graph-to-sequence generation model to convert the graph structure into a text sequence, optimize the generated text sequence at the chapter level, insert the corresponding voice records and image evidence into the corresponding positions according to the semantic association relationship, use the text polishing model to optimize the language of the inserted document content, and generate R&D document data that integrates multimodal information.
8. A research and development document processing device, characterized in that: The device comprises: A multimodal parsing module is used to establish a multimodal document parsing engine, which includes a document format recognition module, a voice recognition module, an image recognition module and a content extraction module. The R&D document is input into the document format recognition module for format type recognition, the recording of the R&D process is input into the voice recognition module for voice-to-text processing, the experimental pictures and equipment photos are input into the image recognition module for feature extraction and scene recognition, and the content extraction module is used to extract text content, voice text, and image annotations to generate raw data containing multimodal information; A cross-modal processing module, used to perform cross-modal preprocessing on the original data, divide the original data into text blocks, voice blocks and image blocks according to information type, perform word segmentation and part-of-speech tagging on the text blocks, construct text semantic vectors, perform speaker recognition and voice segmentation on the voice blocks, extract voice feature vectors, perform target detection and scene classification on the image blocks, generate image feature vectors, input the text semantic vectors, voice feature vectors and image feature vectors into a cross-modal alignment network for semantic matching, and generate standardized document data in a unified feature space; A document standardization module is used to input the standardized document data into a multimodal content understanding model based on deep learning training, identify the R&D elements in the document through the multimodal content understanding model, including R&D goals, technical solutions, innovations, R&D processes, experimental data and R&D results, perform cross-modal semantic association analysis on the identified R&D elements, establish a multimodal relationship map containing text descriptions, voice records and image evidence, reorganize the standardized document data according to a unified document template format based on the multimodal relationship map, and generate R&D document data that integrates multimodal information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the R&D document processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the R&D document processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Standard digital modeling and verification method and system based on artificial intelligence
CN120409658A
A standard digital modeling and verification method and system based on artificial intelligence
CN120409658B
Oral health management system and method based on multi-mode non-language interaction
CN120429831A
Information extraction method and device for unstructured detection text
CN120670589A
Multi-modal data processing method and device and electronic equipment
CN120688021A