Dynamic content identification and conversion method and device, equipment and storage medium
Through optical character recognition, natural language processing and deep learning models, whiteboard content is recognized and converted in real time, which solves the problem of insufficient dynamic content analysis in the existing whiteboard system, and improves work efficiency and standardization of content output.
Patent Information
- Application Number
- CN202510358725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing whiteboard system lacks the automatic analysis and classification function of dynamic content, and cannot effectively learn the discussion process and quickly convert it to standard digital formats, resulting in inefficiency.
Dynamic content recognition and conversion methods are adopted to identify the whiteboard input content in real time through optical character recognition, natural language processing and deep learning models, determine the content type, and convert it into the corresponding target file format, supporting multi-modal information fusion and user preference adjustment.
It realizes automatic classification and analysis of whiteboard content, improves decision-making and execution efficiency, reduces manual processing error rate, and is suitable for scenarios such as education, conferences and business management.
Smart Images

Figure CN120296500A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image data processing, and in particular to a method, device, equipment and storage medium for dynamic content recognition and conversion. Background Art
[0002] Currently, both traditional whiteboards and electronic whiteboards are mainly used for collaboration and visual discussion, and focus on the recognition and storage of static content, but lack effective content analysis and automatic format conversion functions, as well as learning and optimization of dynamic discussion processes. Existing whiteboard systems usually only recognize the final written content and cannot analyze the thinking path of users during meetings or discussions. In addition, current electronic whiteboard tools on the market only support basic hand-drawn conversion and inductive note-taking, lacking automatic classification and analysis functions for the presented and shared content, resulting in output results mostly being image files and PDFs (Portable Document Format), which are essentially also image files, not meeting the actual work execution requirements. Therefore, there is an urgent need for an intelligent whiteboard system that can automatically and intelligently learn the drawing and discussion processes, analyze the content, and quickly convert it into a standard digital format file to improve work efficiency and decision-making accuracy. Summary of the Invention
[0003] In order to solve the technical problems of the lack and inefficiency of the automatic classification and analysis functions of existing whiteboard systems for whiteboard content, the purpose of the present application is to provide a method, device, equipment and storage medium for dynamic content recognition and conversion, and the specific technical solutions adopted are as follows:
[0004] The present application provides a method for dynamic content recognition and conversion, and the method includes:
[0005] Determine the input content in the whiteboard and its content type;
[0006] Use the content type to determine the target conversion file of the input content.
[0007] Preferably, the input content includes at least one of handwritten information, voice information, and text information.
[0008] Preferably, the content type includes at least one of tables, presentation documents, text documents, and digital charts.
[0009] Preferably, the digital charts include at least one of flowcharts, Gantt charts, mind maps, decision trees, data flow diagrams, organization charts, tree structure diagrams, and network topology diagrams.
[0010] Preferably, the determining the input content in the whiteboard and its content type includes:
[0011] Recognize the input content on the whiteboard in real time and determine the matching degree between the input content and each preset file template;
[0012] Determine the file type of the target preset file template with the highest matching degree as the content type of the input content.
[0013] Preferably, the determining the target conversion file of the input content by using the content type includes:
[0014] Map the input content to the target preset file template to generate a target conversion file in the corresponding document format.
[0015] Preferably, after the recognizing the input content on the whiteboard in real time and determining the matching degree between the input content and each preset file template, it further includes:
[0016] Determine the user information of the current user using the whiteboard;
[0017] Use the preference information in the user information to correct the matching degree between the input content and each preset file template.
[0018] Preferably, the determining the input content on the whiteboard and its content type includes:
[0019] Determine the content time series data of the input content on the whiteboard;
[0020] Use the content time series data to adjust and determine the content type of the input content in real time.
[0021] Preferably, the input content includes hand-drawn information; the determining the input content on the whiteboard and its content type includes:
[0022] Use optical character recognition to determine the text information in the hand-drawn information;
[0023] Use natural language processing to determine the content type corresponding to the text information.
[0024] Preferably, the input content includes hand-drawn information; the determining the input content on the whiteboard and its content type includes:
[0025] Use image recognition to determine the graphic information in the hand-drawn information;
[0026] Use a deep learning model to determine the content type corresponding to the graphic information.
[0027] Preferably, the using a deep learning model to determine the content type corresponding to the graphic information includes:
[0028] Use the Mask R-CNN model to segment the graphic information and use the YOLO algorithm to determine the geometric graphic structure in the graphic information;
[0029] Determine the content type corresponding to the graphic information by using each geometric graphic structure.
[0030] Preferably, the input content includes multimodal information arbitrarily composed of hand-drawn information, voice information, and text information; determining the input content and its content type in the whiteboard includes:
[0031] Determine the text data, voice data, and geometric graphic data in the hand-drawn information, voice information, and text information;
[0032] Use the text data, voice data, geometric graphic data, and their respective modal weights to combine and determine the content type corresponding to the multimodal information.
[0033] This application also provides a dynamic content recognition and conversion device, which is used to implement the dynamic content recognition and conversion method described in any one of the above; the device includes:
[0034] A content recognition module, which is used to determine the input content and its content type in the whiteboard;
[0035] A data conversion module, which is used to determine the target conversion file of the input content by using the content type.
[0036] This application also provides a dynamic content recognition and conversion device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the dynamic content recognition and conversion method described in any one of the above are implemented.
[0037] This application also provides a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the dynamic content recognition and conversion method described in any one of the above are implemented.
[0038] This application also provides a computer program product, which when running on a computer causes the computer to execute the dynamic content recognition and conversion method described in any one of the above.
[0039] This application has the following beneficial effects:
[0040] The present application proposes a dynamic content recognition and conversion method, which supports multiple input contents and modes (such as electronic whiteboards, traditional whiteboards, and voice input). In addition to being able to gradually learn the input content and optimize content recognition and judgment during the process of discussing and drawing content, it can also automatically analyze and classify the content, quickly convert the input content into various digital documents that match the corresponding formats of the input content, and is suitable for scenarios such as education, meetings, business management, and process design. It significantly improves the decision-making and execution efficiency, automatically organizes relatively messy input content, improves the work efficiency of standardizing content output, reduces the manual processing error rate and the tediousness of outputting corresponding format files, and is applicable to a variety of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 The flowchart of a dynamic content recognition and conversion method shown in an exemplary embodiment of the present application Figure 1 ;
[0043] Figure 2 The flowchart of a dynamic content recognition and conversion method shown in another exemplary embodiment of the present application Figure 2 ;
[0044] Figure 3 The flowchart of a dynamic content recognition and conversion method shown in another exemplary embodiment of the present application Figure 3 ;
[0045] Figure 4 The flowchart of a dynamic content recognition and conversion method shown in another exemplary embodiment of the present application Figure 4 ;
[0046] Figure 5 The flowchart of a dynamic content recognition and conversion method shown in another exemplary embodiment of the present application Figure 5 ;
[0047] Figure 6 The structural schematic diagram of the hardware operating environment of the dynamic content recognition and conversion device related to the solution of the embodiment of the present application;
[0048] Figure 7 The framework structural schematic diagram of the dynamic content recognition and conversion device related to the solution of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To further elaborate on the technical means and effects adopted by this application to achieve the intended application purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a dynamic content recognition and conversion method proposed according to this application, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0051] The following specifically describes the specific solution of a dynamic content recognition and conversion method provided by this application in conjunction with the accompanying drawings.
[0052] For a dynamic content recognition and conversion method provided by this application, in one embodiment, please refer to Figure 1 , Figure 1 which is a flow chart of a dynamic content recognition and conversion method shown in an exemplary embodiment of this application. Figure 1 .
[0053] The method is applied to a whiteboard and includes:
[0054] Step S101, determining the input content in the whiteboard and its content type;
[0055] Step S102, using the content type to determine the target conversion file of the input content.
[0056] In this embodiment, the whiteboard can be an electronic whiteboard or a traditional whiteboard.
[0057] For an electronic whiteboard, the input content entered into the whiteboard can be collected by means of pen strokes on its touch display screen, or the surrounding voice signals can be collected through its voice collection module, generally a microphone. It can also obtain electronic information such as electronic files through data transmission methods such as USB and wireless.
[0058] For a traditional whiteboard, the written, drawn text symbols and graphics on the traditional whiteboard can be obtained by setting up a camera to collect images. The way of setting up the camera can be to set it at a position slightly above the front of the traditional whiteboard or other positions that can obtain a larger field of view and at a preset distance from the traditional whiteboard.
[0059] The input content in the whiteboard can include at least one or any combination of hand-drawn information, voice information, and text information.
[0060] The hand-drawn information therein includes hand-drawn characters such as letters, numbers, etc., as well as line graphic structures such as arrows, straight lines, curves, boxes, circles, etc. Of course, it can also include any other symbols and graphics that can be presented in a hand-drawn manner.
[0061] For the content types of the input content, it includes at least one of tables, presentation documents, text documents, and digital charts.
[0062] Among them, digital charts include at least one of charts such as flowcharts, Gantt charts, mind maps, decision tree diagrams, tree structure diagrams, data flow diagrams, organization charts, tree structure diagrams, network topology diagrams, timeline diagrams, time plans, calendars, scalable vector graphics, etc.
[0063] In addition, the content types can also be classified in the following ways:
[0064] Type 1: Creativity and thinking organization, including:
[0065] Mindmap: Suitable for creative divergence and brainstorming results.
[0066] Use Case Diagrams: Suitable for requirements analysis and system interaction design.
[0067] Type 2: Project management and time planning, including:
[0068] Gantt Chart: Suitable for schedule planning and project management.
[0069] Timeline: Suitable for project time arrangement and progress monitoring.
[0070] Auto-Schedule Calendar: Suitable for time management and task planning.
[0071] Type 3: Business and process management, including:
[0072] Visio Flowchart: Suitable for process design and business optimization.
[0073] BPMN (Business Process Modeling Notation) Process Script: Suitable for business process modeling and automated execution.
[0074] Visual Workflow: Suitable for operation processes and visual guidance.
[0075] Type 4: Data analysis and decision support, including:
[0076] Excel spreadsheet: suitable for data analysis and report management.
[0077] Decision Trees: suitable for decision analysis and process selection.
[0078] Data Flow Diagram: suitable for data flow and process analysis.
[0079] Type 5: Structure and network planning, including:
[0080] Organization chart: suitable for organizational structure planning and management.
[0081] Architecture Diagrams: suitable for system design and structural analysis.
[0082] Network Topology Diagrams: suitable for network structure planning and analysis.
[0083] Tree Structure Diagrams: suitable for hierarchical visualization and analysis.
[0084] Type 6: Presentation and archiving, including:
[0085] PowerPoint presentation: suitable for presentations and proposal reports.
[0086] PDF document: suitable for standardized document archiving and sharing.
[0087] SVG (Scalable Vector Graphics) graphics: suitable for visual presentation and editing.
[0088] For the recognition and classification of the input content, the existing mature optical character recognition technology (OCR) can be used to recognize the text, numbers and other character symbols, and the natural language processing technology (NLP) can be used to recognize the text information composed of words. In addition, the existing neural network models for recognizing graphic information can be used to extract the graphic features in the input content and distinguish the graphic structures. For the speech information in the input content, some speech recognition models can be used to recognize it and convert the speech information into text information or other information after noise reduction.
[0089] After determining the content type of the input content, a preset file template matching the content type can be retrieved, and the input content can be directly converted into graphic materials in the preset file template, or the content can be converted into characters and schemas in a standard format and then filled into the corresponding positions in the preset file template, so as to generate a target conversion file that meets the corresponding file specification standards. For example, after determining that the hand-drawn information is an arrow in a flowchart, the hand-drawn information can be converted into a standard arrow material in the flowchart file template; for another example, after determining that the hand-drawn information is a hand-drawn table, the data in the hand-drawn table can be converted into a standard font format and then filled into the table file template, or the data and table frame lines in the hand-drawn table can be simultaneously converted into the frame lines and data in a standard spreadsheet. If the input content is voice information, it can be at least converted into various formats of text documents or presentation documents. In addition, the content type corresponding to the hand-drawn information can be further determined through the voice information, so as to convert the hand-drawn information into a target conversion file that better meets the actual needs of the user. If the input content itself is text information, it can be converted into text documents or presentation documents in other formats.
[0090] The embodiment of the present application proposes a dynamic content recognition and conversion method, which supports various input contents and modes (such as electronic whiteboard, traditional whiteboard and voice input). In addition to being able to gradually learn the input content and optimize content recognition and judgment during the process of discussing and drawing content, it can also automatically analyze and classify the content, quickly convert the input content into various digital documents matching the corresponding formats of the input content, and is suitable for scenarios such as education, meetings, business management and process design. It significantly improves the decision-making and execution efficiency, automatically sorts out relatively messy input content, improves the work efficiency of standardizing content output, reduces the manual processing error rate and the tediousness of outputting files in the corresponding format, and is applicable to various application scenarios.
[0091] For a dynamic content recognition and conversion method provided by the present application, in one embodiment, please refer to Figure 2 , Figure 2 is the flow of a dynamic content recognition and conversion method shown in another exemplary embodiment of the present application Figure 2 .
[0092] In this embodiment, the method includes:
[0093] Step S201, real-time recognize the input content on the whiteboard and determine the matching degree between the input content and each preset file template;
[0094] Step S202, determine the file type of the target preset file template with the highest matching degree as the content type of the input content;
[0095] Step S203, use the content type to determine the target conversion file of the input content.
[0096] In this embodiment, the input content input to the whiteboard can be obtained and recognized dynamically and in real time. By means of deep learning and neural network models to recognize and analyze the input content, the input content is input into the corresponding neural network model, such as the Transformer model. The data training set corresponding to this neural network model can include a large number of various file templates, including various tables, presentation documents, text documents, and digital charts mentioned in the above embodiments. Thus, the trained neural network model determines the matching degree, that is, the similarity, between the input content and each preset file template, determines the target preset file template with the highest matching degree as the content type of the input content, and then maps and converts the input content into a target conversion file according to the target preset file template.
[0097] In a specific embodiment, the input content includes handwritten information; the step S101 includes:
[0098] Using optical character recognition to determine the text information in the handwritten information;
[0099] Using natural language processing to determine the content type corresponding to the text information.
[0100] Through optical character recognition technology, handwritten characters, symbols, or digital printed characters, symbols can be recognized and converted into editable text, such as in the word document format, or non-editable PDF document format.
[0101] Optical character recognition technology (OCR) can also process images through a convolutional neural network (CNN) to extract text information.
[0102] The process generally includes: image denoising → feature extraction → character segmentation → text recognition → result output.
[0103] Specific key models or tools that can be used include: Tesseract OCR software or other deep learning models (such as CRNN).
[0104] Among them, optical character recognition mainly relies on a convolutional neural network (CNN) for image feature extraction and character recognition.
[0105] Meanwhile, Tesseract OCR software can also be used as a basic tool, and combined with deep learning models (such as CRNN) to improve the accuracy of handwritten recognition. For non-standard fonts, training is carried out through enhanced datasets (rotation, blurring, etc.). Convolution operations are used to extract local features in the image, such as the edge shapes of letters and numbers. The ReLU activation function can also be used during model training to improve the extraction effect of non-linear features. The CTC (Connectionist Temporal Classification) loss function is used to ensure that the order of the output characters matches the input.
[0106] In addition, dynamic font adaptation can be performed: Based on multi-modal learning technology, the recognition accuracy of handwritten and non-standard fonts is enhanced, and the convolution kernel size is dynamically adjusted through a feature adaptation layer.
[0107] More specifically, for the deep learning model used, CRNN (Convolutional Recurrent Neural Network) is used for illustration. CRNN is used for character detection and sequence prediction:
[0108] The convolutional layer therein: is used to extract image features.
[0109] The RNN (Recurrent Neural Network) layer therein: processes sequence dependencies and generates the character order.
[0110] The CTC loss function: ensures that the output character sequence is consistent with the character order in the image.
[0111] OCR technology can be used to extract text from the content displayed on an electronic whiteboard or obtained through data transmission, or the content photographed from a traditional whiteboard. This is applicable to the digitization of meeting minutes or classroom notes. Examples of application scenarios are as follows:
[0112] Scenario 1: Automatic archiving of meeting notes. Using high-precision OCR to extract the photographed whiteboard text into editable text to achieve the digitization of meeting notes.
[0113] Scenario 2: Digital conversion and archiving of documents.
[0114] After recognizing the text information in the handwritten information, natural language processing (NLP) technology can be further applied to analyze the text semantics of the whole or part of the text information through NLP technology, extract keywords, themes and classifications, automatically identify the phased content of meetings or teaching discussions, and classify the content into theme categories such as brainstorming, strategic planning, business management, process design, etc. according to the context. The input content can also be classified and archived in this way of theme classification. At the same time, content tagging and sentiment analysis can also be provided to improve the classification accuracy.
[0115] Based on pre-trained language models such as GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representations from Transformers), analyze semantics and perform content classification.
[0116] The process generally includes: text tokenization → keyword extraction → theme / content classification → semantic modeling → recommended output template.
[0117] Purpose: Used for automatic classification and template matching to generate appropriate document formats.
[0118] The multi-head attention mechanism in the model helps the model capture long-range dependencies in the input text, such as the association between sentence themes and details. The classification loss function ensures that the model can correctly distinguish different types of content.
[0119] The sentiment analysis therein combines session topic recognition: analyzes the sentiment tendency of semantics through a language sentiment model (such as RoBERTa), and at the same time recognizes the session topic and combines the attention mechanism of the Transformer structure of the model itself to achieve the function of automatic text summarization.
[0120] Use the pre-trained BERT model for semantic classification and extract the theme of the content in combination with sentence vectors (Sentence Embedding). Automatic summary generation and template recommendation are based on the Transformers structure.
[0121] The entire NLP mentioned above can be used for automatic classification and analysis of whiteboard text content. For example, classify meeting discussions as "brainstorming" or "progress planning", and recommend appropriate document templates, such as word documents, PDF documents, PPT presentation documents, electronic notes or other chart documents, etc.
[0122] Here are two actual application scenarios:
[0123] Application Scenario 1: Meeting minutes generation, parse the text discussed and drawn during the meeting, extract key points and automatically generate a summary.
[0124] Application Scenario 2: Course induction in the education scenario, semantic classification of classroom dialogue data and hand-drawn data, and generation of course notes.
[0125] In another specific embodiment, the input content includes hand-drawn information; step S101 includes:
[0126] Use image recognition to determine the graphic information in the hand-drawn information;
[0127] Use a deep learning model to determine the content type corresponding to the graphic information.
[0128] The use of a deep learning model to determine the content type corresponding to the graphic information specifically includes:
[0129] Use the Mask R-CNN model to segment the graphic information and use the YOLO algorithm to determine the geometric graphic structure in the graphic information;
[0130] Use each geometric graphic structure to determine the content type corresponding to the graphic information.
[0131] In this embodiment, a convolutional neural network (CNN) model can be applied to recognize graphic structures (such as arrows, lines, block diagrams, etc.), and is applied to recognize hand-drawn graphics or whiteboard content.
[0132] The process generally includes: image segmentation → feature detection → pattern matching → conversion to a standardized data format.
[0133] Purpose: Convert a hand-drawn flowchart into a chart format such as Visio or SVG (Scalable Vector Graphics).
[0134] Among them, common semantic segmentation models such as U-Net, the core of which is to perform downsampling and upsampling on images.
[0135] Among them, Hu moment invariant features can also be applied to detect specific graphics (such as arrows or block diagrams), and these shapes can be converted into standard symbols in chart tools such as Visio or SVG.
[0136] OpenCV technology and the YOLO algorithm (a real-time object detection algorithm) can also be used to detect geometric graphics (such as arrows, lines, rectangles, circles, polygons, etc.) on the whiteboard. Combine Mask R-CNN for image segmentation to separate and label each graphic.
[0137] The loss function in the CNN model is used for the training of the semantic segmentation model to distinguish the graphic area from the background. Reinforcement learning for graphic detection can also be carried out: applying reinforcement learning to the detection of process graphics (arrows, block diagrams, etc.), and continuously optimizing the detection results through environmental feedback. In addition, the FPN (Feature Pyramid Network) can be used to process images and perform detection by combining multi-level features.
[0138] The core model for realizing the above image (graphic) recognition can be selected as: Mask R-CNN, a specific CNN model.
[0139] Its functions and uses include: detecting hand-drawn arrows, rectangles, and circles in the whiteboard image. Using ResNet (Residual Network) as the feature extractor to generate region proposals.
[0140] Data labeling and training of the model: Using the LabelImg tool to label data and convert the format to the COCO (Common Objects in Context) format.
[0141] Specific model deployment schemes and key points can include:
[0142] Model service-ization: Using TorchServe to deploy Mask R-CNN and write a custom model processor to return the detected graphic types and coordinates.
[0143] Edge deployment: Deploying a lightweight model (such as the MobileNet version of Mask R-CNN) to the device side to process real-time whiteboard images. The image recognition technology can convert hand-drawn flowcharts and structure diagrams into a standardized format (such as a Visio flowchart).
[0144] Here are two actual application scenarios:
[0145] Application scenario 1: Automatic conversion of design sketches, converting hand-drawn design drawings into vector graphics, and adapting to tools such as Visio or CAD.
[0146] Application scenario 2: Digitalization of process management, identifying flowcharts and generating standardized business process scripts.
[0147] In another embodiment, for the case where the input content includes voice information, the voice input is transcribed into text through speech recognition technology and automatically classified by a semantic analysis model.
[0148] Using a deep learning voice model (such as Wav2Vec or DeepSpeech) to transcribe the voice into text.
[0149] The process of speech recognition generally includes: audio denoising → feature extraction → speech-to-text → semantic classification.
[0150] Based on the use of existing speech models, the following key points can also be included:
[0151] Use Mel Frequency Cepstrum Coefficient (MFCC) to represent speech signals and train the speech model based on the CTC (Connectionist Temporal Classification) loss function.
[0152] MFCC is used to extract the core features of speech audio, and these features can be used to train the speech recognition model. The CTC loss function ensures that the output text sequence is consistent with the actual speech content.
[0153] During the process of speech recognition, real-time speech transcription optimization can also be carried out: use an attention-based acoustic model (such as LAS, Listen-Attend-Spell) to process the speech stream, and speech segmentation and speaker recognition can also be performed. For example, apply the x-vector technology for speaker segmentation to distinguish the content of multiple speakers.
[0154] Taking the DeepSpeech model as an example, the process of speech recognition implementation includes the following main links and tools to be used:
[0155] Use DeepSpeech for speech-to-text transcription, extract speech features through MFCC, and use BiLSTM (Bidirectional Long Short-Term Memory Network) to capture the feature dependencies of time steps.
[0156] Here are two actual application scenarios:
[0157] Application scenario 1: Multilingual transcription in remote collaboration, supporting the generation of multilingual meeting records.
[0158] Application scenario 2: Automatic generation of diagnostic records, automatically transcribing the speech descriptions in the medical scenario into structured reports.
[0159] In another specific embodiment, the input content includes multimodal information composed of any of hand-drawn information, speech information, and text information; the step S101 includes:
[0160] Determine the text data, speech data, and geometric graphic data in the hand-drawn information, speech information, and text information;
[0161] Utilize the text data, speech data, geometric graphic data, and their respective modal weights to combine and determine the content type corresponding to the multimodal information.
[0162] In this embodiment, the input content for the whiteboard may not be limited to a single type of hand-drawn information, voice information, or text information, but rather multimodal information composed of a combination of two or three or more types of information (the content type can be referred to as mixed content). It is necessary to perform multimodal data fusion on the multimodal information, mainly by combining text, image, and voice data to construct a unified multimodal representation model.
[0163] The process may include: data preprocessing → multimodal feature extraction → fusion modeling → result output.
[0164] The uses of the multimodal representation (learning) model: support the processing of multiple types of inputs in complex scenarios (such as voice + image + text).
[0165] The fusion of multiple models: Combine image, text, and voice data to generate a multimodal representation.
[0166] Specifically, TensorFlow or PyTorch can be used to construct a multimodal learning model, combining image, voice, and text features. Dynamically adjust the modal weights to improve the fusion effect and enhance the ability to understand mixed content.
[0167] One important ability of the multimodal learning model is adaptive weight fusion, that is, it can dynamically adjust the weights of each data according to the importance of the input data modality, here referring to the modal weights of hand-drawn information, voice information, or text information respectively, or the weights between voice data, graphic data, and text data.
[0168] Specifically for the content fusion strategy, BERT (Bidirectional Encoder Representations from Transformers) can be used to extract text embeddings, ResNet (Residual Neural Network) to extract image embeddings, and finally a multimodal Transformer model is used for data fusion.
[0169] Multimodal technology can fuse voice, image, and text data to generate a unified data representation and achieve more advanced content analysis. Here is an exemplary application scenario:
[0170] Multimodal classroom note generation: Fuse voice, whiteboard content, and note pictures to generate complete digital notes.
[0171] For a dynamic content recognition and conversion method provided by this application, in one embodiment, please refer to Figure 3 , Figure 3 which shows the process of a dynamic content recognition and conversion method in another exemplary embodiment of this application Figure 3 .
[0172] In this embodiment, the method includes:
[0173] Step S301: Real-time identify the input content on the whiteboard and determine the matching degree between the input content and each preset file template;
[0174] Step S302: Determine the file type of the target preset file template with the highest matching degree as the content type of the input content;
[0175] Step S302: Map the input content to the target preset file template to generate a target conversion file in the corresponding document format.
[0176] Based on the above embodiments, after identifying the input content and the content type of the input content through deep learning and determining the target preset file template that matches it, the identified and classified input content can be mapped to the target preset file template to generate the corresponding digital document format, that is, some irregular and discrete input content is converted and filled into each standardized material and element in the target preset file template to form the corresponding standardized and structured target conversion file.
[0177] This embodiment supports multiple document format outputs, including but not limited to:
[0178] Mindmap: Used for creative organization and brainstorming.
[0179] Gantt Chart: Suitable for project progress management.
[0180] Visio flowcharts and organization charts: Used for business process optimization and structural analysis.
[0181] Excel spreadsheet: Suitable for data analysis and financial management.
[0182] BPMN (Business Process Modeling Notation) process script: Used for business process modeling and automated execution.
[0183] For a dynamic content recognition and conversion method provided in this application, in one embodiment, please refer to Figure 4 , Figure 4 which is the flow of a dynamic content recognition and conversion method shown in another exemplary embodiment of this application Figure 4 .
[0184] The method includes:
[0185] Step S401, identify the input content on the whiteboard in real time, and determine the matching degree between the input content and each preset file template:
[0186] Step S402, determine the user information of the current user using the whiteboard;
[0187] Step S403, use the preference information in the user information to correct the matching degree between the input content and each preset file template.
[0188] Based on the above embodiments, after initially obtaining the matching degree between the input content and each preset file template, the matching degree between the input content and each preset file template can be optimized and adjusted by analyzing the relevant information of the user, with the aim of making the finally matched target preset file template and the generated target conversion file meet the actual needs and usage habits of the user.
[0189] This embodiment is based on a reinforcement learning model. According to the user's preference information, such as which type of file format the user tends to or is accustomed to using, the user's preference information can be determined by analyzing the user's historical usage behavior. For example, based on the target template file independently selected by the user after drawing on the whiteboard, or setting the user's own preferences in the early stage of using the whiteboard, the whiteboard system can provide corresponding option content, and the user can freely set the whiteboard system to automatically generate the corresponding target conversion file after using the whiteboard. For example, some users like to generate PDF document files from their hand-drawn content, and some users like to generate editable word files from their voice teaching content. Based on the user's preference information, the matching degree of the preferred preset file template can be adjusted to the highest or increased by a certain degree, making it easier to recommend the preset file template preferred by the user.
[0190] In addition, the user information of the current user can be determined by the current account logged in by the user through the electronic whiteboard system, or the identity information or identity characteristics of the user can be determined by face recognition through the built-in or external camera of the electronic whiteboard. For traditional whiteboards, the current user and their user information can be determined by external cameras or pickup devices through face recognition or voiceprint recognition.
[0191] Regarding how to correct the matching degree between the input content and each preset file template based on preference information through deep learning, the following implementation process can be included:
[0192] Basic process: Content feature extraction → Similarity calculation → Template sorting → Automatic matching.
[0193] Application scenario: Automatically select the best output format (such as Gantt chart or mind map).
[0194] Reinforcement learning-based recommendation algorithms, such as using the Q-learning method to optimize the recommendation of preset file templates. First, cosine similarity can be used to measure the matching degree between the input content and the preset file templates, that is, calculate the similarity between the content and the template based on cosine similarity. Use PyTorch to implement the comparison of template embedding vectors. The reward signal of Q-learning can be determined by the user's historical selection of file templates or the preference for freely preset output file templates, continuously optimizing the accuracy of template recommendations.
[0195] Among them, cosine similarity is used to measure the matching degree between the input content and the preset file templates, ensuring the relevance of the recommendation results. Further, based on the template optimization of contrastive learning, contrastive loss is used for template optimization and similarity-based recommendation. In addition, Faiss can be used for fast retrieval of similar file templates.
[0196] In the actual application process, the whiteboard's template recommendation system recommends appropriate output formats (such as Gantt charts or mind maps) for users according to the input content features or gives suggestions on intelligent output formats.
[0197] For a dynamic content recognition and conversion method provided by this application, in a specific embodiment, please refer to Figure 5 , Figure 5 which is the flowchart of a dynamic content recognition and conversion method shown in another exemplary embodiment of this application. Figure 5 .
[0198] The dynamic content recognition and conversion method includes:
[0199] Step S501, determining the content time series data of the input content in the whiteboard;
[0200] Step S502, using the content time series data to adjust and determine the content type of the input content in real time;
[0201] Step S503, using the content type to determine the target conversion file of the input content.
[0202] In this embodiment, considering that the input content may be modified, deleted, added, etc. at any time, it is necessary to dynamically and real-time adjust and determine the content type of the input content according to the changes in the input content. For this, the process and results of recognizing the input content can be continuously learned and optimized through deep learning. Specifically, taking the handwritten information as an example, the content time-series data of the handwritten information is obtained, that is, the handwritten data at each acquisition moment. Based on the content time-series data, the order of whiteboard writing and drawing is analyzed through a deep learning model, and the stroke changes, writing directions, and content structures are recorded, so as to determine the content type of the input content in real time. As the handwritten data increases, the content type can be recognized more accurately, and the true intention of the user's handwritten and the target conversion file required can be determined.
[0203] For the handwritten information, the corresponding dynamic recognition process is: handwriting capture → time series analysis → model learning and recognition → content classification and structure optimization.
[0204] More specifically, the data features in the handwritten information are extracted. Through the stroke vectorization technology, the handwritten strokes are converted into digital sequences, that is, the content time-series data. Furthermore, the convolutional neural network (CNN) and the recurrent neural network (RNN) are applied to classify the digitized handwriting, so that the drawing and writing order can be inferred based on the content time-series data, and the flow chart and structured information can be determined. Finally, the preset file template corresponding to the determined content type can be retrieved to generate the target conversion file corresponding to the handwritten information. For example, the manually drawn flow chart and structure diagram can be converted into electronic flow charts, structure diagrams and other charts in the corresponding standardized data format, such as charts or graphic structures in Visio or CAD format.
[0205] Corresponding application scenarios, for example, may include: identifying key stages of a meeting discussion, such as brainstorming, planning, decision-making, etc.
[0206] The embodiment of the present application also proposes a dynamic content recognition and conversion device. The dynamic content recognition and conversion device can be a display and data processing device such as an electronic whiteboard, a conference tablet, a smart TV, a computer, etc.
[0207] As Figure 6 shown, Figure 6 is a schematic structural diagram of the hardware operating environment of the dynamic content recognition and conversion device involved in the embodiment of the present application.
[0208] As Figure 6As shown in the figure, the dynamic content recognition and conversion device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, a communication bus 1002, and a camera 1006. Among them, the communication bus 1002 is used to implement the connection and communication between these components. The user interface 1003 may include a display and an input unit such as a control panel. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. The memory 1005, as a computer storage medium, may include a dynamic content recognition and conversion program. Among them, the camera 1006 may be an RGB (red, green, blue) camera.
[0209] Those skilled in the art can understand that Figure 6 the hardware structure shown in the figure does not constitute a limitation on the device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0210] Continuing to refer to Figure 6 , Figure 6 the memory 1005, as a computer-readable storage medium, may include an operating device, a user interface module, a network communication module, and a dynamic content recognition and conversion program.
[0211] In Figure 6 , the network communication module is mainly used to connect to the server and can communicate with the server for data; while the processor 1001 can call the dynamic content recognition and conversion program stored in the memory 1005 and execute the steps in each of the above embodiments.
[0212] Based on the hardware structure of the above dynamic content recognition and conversion device, each embodiment for implementing the dynamic content recognition and conversion method of the present application is realized.
[0213] In addition, the present application also provides a dynamic content recognition and conversion device. Please refer to Figure 7 , the dynamic content recognition and conversion device includes:
[0214] a content recognition module A10 for determining the input content in the whiteboard and its content type;
[0215] a data conversion module A20 for determining the target conversion file of the input content by using the content type.
[0216] Furthermore, the content recognition module A10 is further configured to:
[0217] Real-time recognize the input content in the whiteboard and determine the matching degree between the input content and each preset file template;
[0218] Determine the file type of the target preset file template with the highest matching degree as the content type of the input content.
[0219] Furthermore, the data conversion module A20 is further configured to:
[0220] Map the input content to the target preset file template to generate a target conversion file in the corresponding document format.
[0221] Furthermore, the content recognition module A10 is further configured to:
[0222] Determine the user information of the current user using the whiteboard;
[0223] Utilize the preference information in the user information to correct the matching degree between the input content and each preset file template.
[0224] Furthermore, the content recognition module A10 is further configured to:
[0225] Determine the content time series data of the input content in the whiteboard;
[0226] Utilize the content time series data to adjust and determine the content type of the input content in real time.
[0227] Furthermore, the content recognition module A10 is further configured to:
[0228] Utilize optical character recognition to determine the text information in the handwritten information;
[0229] Utilize natural language processing to determine the content type corresponding to the text information.
[0230] Furthermore, the content recognition module A10 is further configured to:
[0231] Utilize image recognition to determine the graphic information in the handwritten information;
[0232] Utilize a deep learning model to determine the content type corresponding to the graphic information.
[0233] Furthermore, the content recognition module A10 is further configured to:
[0234] Utilize the Mask R-CNN model to segment the graphic information and use the YOLO algorithm to determine the geometric graphic structure in the graphic information;
[0235] Utilize each geometric graphic structure to determine the content type corresponding to the graphic information.
[0236] Furthermore, the content recognition module A10 is further configured to:
[0237] Determine the text data, voice data, and geometric graphic data in the hand-drawn information, voice information, and text information;
[0238] Use the text data, voice data, geometric graphic data, and their respective modality weights to combine and determine the content type corresponding to the multi-modal information.
[0239] The specific implementation manner of the dynamic content recognition and conversion device of the present application is basically the same as that of the above-mentioned embodiments of the dynamic content recognition and conversion method, and will not be elaborated here.
[0240] In addition, the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium of the present application, and the computer program can be a dynamic content recognition and conversion program. When the dynamic content recognition and conversion program is executed by a processor, the steps of the above-mentioned dynamic content recognition and conversion method are implemented.
[0241] Wherein, the method implemented when the dynamic content recognition and conversion program is executed can refer to the various embodiments of the dynamic content recognition and conversion method of the present application, and will not be elaborated here.
[0242] In addition, the present application also provides a computer program product. The computer program product includes computer program code. When the computer program code runs on a computer, the computer is caused to execute the dynamic content recognition and conversion method in the above various implementation manners.
[0243] It should be noted that: the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0244] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.
[0245] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0246] The above are only the preferred embodiments of the present application, and do not limit the protection scope of the present application. Any equivalent structure / method transformation made under the application concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the protection scope of the present application.
Claims
1. A method for dynamic content recognition and conversion, characterized in that Applied to a whiteboard, the method includes: Determine the input content in the whiteboard and its content type; Use the content type to determine the target conversion file of the input content.
2. The dynamic content recognition and conversion method according to claim 1, wherein The input content includes at least one of handwritten information, voice information, and text information.
3. The dynamic content recognition and conversion method according to claim 1, characterized in that The content type includes at least one of a table, a presentation, a text document, and a digital chart.
4. The dynamic content recognition and conversion method according to claim 3, wherein The digital chart includes at least one of a flowchart, a Gantt chart, a mind map, a decision tree diagram, a data flow diagram, an organization chart, a tree structure diagram, and a network topology diagram.
5. The dynamic content recognition and conversion method according to claim 1, wherein The determining the input content in the whiteboard and its content type includes: Real-time identify the input content in the whiteboard and determine the matching degree between the input content and each preset file template; Determine the file type of the target preset file template with the highest matching degree as the content type of the input content.
6. The dynamic content recognition and conversion method according to claim 5, wherein The using the content type to determine the target conversion file of the input content includes: Map the input content to the target preset file template to generate a target conversion file in the corresponding document format.
7. The dynamic content recognition and conversion method according to claim 5, characterized in that After the real-time identifying the input content in the whiteboard and determining the matching degree between the input content and each preset file template, it further includes: Determine the user information of the current user using the whiteboard; Use the preference information in the user information to correct the matching degree between the input content and each preset file template.
8. The dynamic content recognition and conversion method according to claim 1, wherein The determining the input content in the whiteboard and its content type includes: Determine the content time series data of the input content in the whiteboard; Use the content time series data to adjust and determine the content type of the input content in real time.
9. The dynamic content recognition and conversion method according to claim 1, wherein The input content includes handwritten information; the determining the input content in the whiteboard and its content type includes: Use optical character recognition to determine the text information in the handwritten information; Use natural language processing to determine the content type corresponding to the text information.
10. The dynamic content recognition and conversion method according to claim 1, wherein The input content includes handwritten information; the determining the input content in the whiteboard and its content type includes: Use image recognition to determine the graphic information in the handwritten information; Use a deep learning model to determine the content type corresponding to the graphic information.
11. The dynamic content recognition and conversion method according to claim 10, wherein The using the deep learning model to determine the content type corresponding to the graphic information includes: Use the Mask R-CNN model to segment the graphic information and use the YOLO algorithm to determine the geometric graphic structure in the graphic information; Use each geometric graphic structure to determine the content type corresponding to the graphic information.
12. The dynamic content recognition and conversion method according to claim 1, wherein The input content includes multimodal information composed of any of handwritten information, voice information, and text information; the determining the input content in the whiteboard and its content type includes: Determine the text data, voice data, and geometric graphic data in the handwritten information, voice information, and text information; Use the text data, voice data, geometric graphic data, and their respective modal weights to combine and determine the content type corresponding to the multimodal information.
13. A dynamic content recognition and conversion device, characterized in that, It includes: A content recognition module for determining the input content in the whiteboard and its content type; A data conversion module for using the content type to determine the target conversion file of the input content.
14. A dynamic content recognition and conversion device, characterized in that The dynamic content recognition and conversion device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the dynamic content recognition and conversion method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the steps of the dynamic content recognition and conversion method according to any one of claims 1 to 12 are implemented.
16. A computer program product, characterized in that, When it runs on a computer, it causes the computer to execute the dynamic content recognition and conversion method according to any one of claims 1 to 12.
Citation Information
Cited By
PPT style conversion method and system
CN121788662A
A ppt style conversion method and system
CN121788662B