Intelligent pen data acquisition and processing method based on multi-mode sensor
Through the smart pen integrating the pen tip pressure sensor, OCR camera group and microphone, the synchronous acquisition and processing of multimodal data is realized, solving the problem of insufficient integration of existing smart pen functions and providing efficient information collection and organization capabilities.
Patent Information
- Application Number
- CN202510513434.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing smart pen products have insufficient functional integration, which is difficult to meet the needs of synchronized text, images, and voice recording in multiple scenarios. The complex interaction design leads to delayed information collection and low data management efficiency, making it difficult to apply in-depth in fields such as mobile office and academic research.
It adopts a smart pen based on multimodal sensor, integrates a pen tip pressure sensor, an OCR camera group and a microphone, and synchronous acquisition of image and audio data through a dual trigger mechanism, combines the main control chip for data integration and processing, and uses a Bluetooth module to transmit it to the mobile APP in real time.
It realizes synchronous recording of images and audio in multiple scenarios, ensures data timestamp alignment through dynamic time regularization algorithm, removes redundancy, and uses GATT protocol for efficient data transmission, breaking through the single functional limitation of traditional smart pens and providing convenient information collection and organization capabilities.
Smart Images

Figure CN120447760A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart pens, and in particular to a method for collecting and processing smart pen data based on a multimodal sensor. Background Art
[0002] In the context of digital transformation, users' demand for multimodal information recording is growing. Smart pens have become one of the common tools people use for daily information recording and play an important role in digital recording scenarios. However, smart pen products on the current market still have technical shortcomings, such as insufficient functional integration. Most devices only support handwriting or a single multimedia acquisition function, which is difficult to meet the needs of synchronous text, image, and voice recording in scenarios such as meeting minutes and inspiration capture; there are defects in the interactive design, and complex operating processes (such as multi-level menu settings and multi-device switching) lead to delays in the collection of key information and low data management efficiency. Handwritten handwriting and audio and video files need to be exported through different channels, and subsequent sorting is time-consuming and labor-intensive. These technical shortcomings restrict the in-depth application of smart pens in mobile office, academic research and other fields. There is an urgent need to achieve breakthrough improvements in functional integration and user experience through innovative design.
[0003] In order to solve the above problems, a smart pen data acquisition and processing method based on multimodal sensors is proposed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects of the existing technology and propose a smart pen data collection and processing method based on a multimodal sensor. By using this device to work, the problem of inconvenience in timely collection and organization of fragmented information is solved.
[0005] To solve the above problems, the present invention adopts a technical solution: a smart pen data collection and processing method based on a multimodal sensor, comprising the following steps: S1: By squeezing the pen tip pressure sensor and finger pressure sensor, a corresponding trigger signal is obtained. Based on the trigger signal, the OCR camera group is controlled to start, and the content in front of the pen tip is photographed to obtain image data, and the image data is stored; S2: Trigger the recording switch by rotating the pen end, turn on the microphone to record, obtain audio data and store it; S3: The stored image data and audio data are integrated and processed by the main control chip to obtain integrated data and store it; S4: The stored image data, audio data and integrated data are transmitted to the mobile phone APP in real time through the Bluetooth module.
[0006] Furthermore, S1 includes the following steps: S11: When the user writes normally, the pen tip is squeezed to trigger the pen tip pressure sensor, and the pen tip pressure sensor collects the pen tip pressure data of the writing process of the smart collection pen in real time to obtain the smart collection pen tip pressure trigger signal; S12: The tip pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the tip pressure sensor of the pen. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the digital signal of the smart collection pen. S13: The smart collection pen digital signal is filtered to remove noise through a filtering algorithm to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip. The smart collection pen digital control signal is transmitted to the OCR camera group to activate the camera hardware and add an initial image timestamp. S14: The optical image data captured by the OCR camera group is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered, and converted into digital image data through an analog-to-digital converter. The digital image data is finally compressed and encoded by the main control chip and written to the storage medium, and an end image timestamp is added.
[0007] Furthermore, S1 further includes the following steps: S15: When it is inconvenient to squeeze the pen tip when collecting information on the electronic device, the user presses the finger pressure switch to trigger the finger pressure sensor, and the finger pressure sensor collects finger pressure data of the pressing process of the smart collection pen in real time to obtain a finger pressure trigger signal of the smart collection pen; S16: The finger pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the finger pressure sensor. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the digital signal of the smart collection pen. S17: The smart collection pen digital signal is subjected to a filtering algorithm to remove noise to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip. The smart collection pen digital control signal is transmitted to the OCR camera group to activate the camera hardware and add an initial image timestamp. S18: The optical image data captured by the OCR camera group is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered, and converted into digital image data through an analog-to-digital converter. Finally, the main control chip compresses and encodes the digital image data and writes it to the storage medium, adding an end image timestamp.
[0008] Furthermore, S2 includes the following steps: S21: The user rotates the recording switch at the end of the pen, and the mechanical action signal generated by the recording switch is converted into a trigger signal for the smart acquisition laptop through the Hall sensor; S22: The trigger signal of the smart collection laptop is shaped and amplified by the signal conditioning circuit and then transmitted to the main control chip. After receiving the trigger signal of the smart collection laptop, the main control chip generates a recording control signal of the smart collection pen. The recording control signal of the smart collection pen is transmitted to the microphone at the end of the pen to start the recording function and add an initial recording timestamp mark at the same time; S22: The audio sound wave data recorded by the microphone at the end of the pen is converted into an audio analog electrical signal through the electret capacitor. The audio analog electrical signal is converted into digital audio data by the analog-to-digital converter after pre-amplification and filtering, and transmitted to the main control chip through the audio transmission line. The main control chip compresses and encodes the data and writes it to the storage medium, and adds a timestamp to mark the end of recording.
[0009] Furthermore, S18 includes the following steps: S181: The digital image data is processed by grayscale, noise reduction, and binarization to optimize the image quality and obtain pre-processed image data; S182: Preprocessing the image data to extract text features through an image recognition module to obtain graphic information data; S183: The graphic information data is processed through character correction and semantic verification calculation to obtain intelligent graphic data.
[0010] Furthermore, S22 includes the following steps: S221: Converting the digital audio data from the time domain to the frequency domain through Fourier transform to obtain noise audio data; S222: The digital audio data is processed by an adaptive filtering algorithm, and filtering is performed according to the determined noise audio data to remove noise interference and obtain filtered audio data; S223: The filtered audio data is compared with a set energy threshold to detect blank sound segments in the audio, and the blank sound segments are deleted from the filtered audio data to obtain simplified audio data.
[0011] Furthermore, S22 further includes the following steps: S224: Simplify the audio data by extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients to obtain audio feature data; S225: The audio feature data is processed and analyzed by matching with the acoustic model and the language model to obtain transcribed text data; S226: The transcribed text data is processed through character correction and semantic verification calculations to obtain intelligent audio-text data.
[0012] Furthermore, S3 includes the following steps: S31: Integrate the intelligent graphic data and the intelligent audio data to obtain text data set data; S32: The text data set is analyzed and processed using the cosine similarity algorithm and the edit distance algorithm to calculate the similarity of the text information in the data set, find duplicate text content, and obtain redundant text data; S33: Delete redundant text data from the text data set to obtain intelligent text set data.
[0013] Furthermore, S31 includes the following steps: S311: Collecting the initial image timestamp, the final image timestamp, the initial recording timestamp, and the final recording timestamp in the intelligent image and text data and the intelligent audio and text data; S312: Comparing the intelligent text set data with the intelligent audio and text data and segmenting the time series, segmenting the time series of the intelligent audio and text data and the intelligent text set data; S313: Use the dynamic time warping algorithm to match and align the segmented intelligent audio and text data and the intelligent text set data, adjust the timestamps, and obtain the intelligent synchronization set data.
[0014] Furthermore, S4 includes the following steps: S41: The smart collection pen completes pairing and establishes a connection with the mobile phone app through the Bluetooth module. At the same time, the smart collection device organizes and encapsulates the smart synchronization set data, smart text set data, smart image and text data, and smart audio and text data into binary protocol. By adding a data type identification header, a data length field, and a CRC check code, different types of data are packaged according to the frame structure to prepare for data transmission. S42: The intelligent data collection device sends the encapsulated data to the mobile phone APP in real time through the Bluetooth module. After receiving the data, the mobile phone APP receives the data by subscribing to the characteristic value based on the Bluetooth GATT protocol, calls the corresponding parser according to the predefined frame header identifier, and reorganizes the data into JSON format for processing.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a smart pen data acquisition and processing method based on multimodal sensors. The smart pen integrates modules such as a pen tip pressure sensor, an OCR camera group, and a microphone. The user triggers image acquisition by squeezing the pen tip or pressing the switch with his finger, and simultaneously records the initial and end timestamps. The image is pre-processed, OCR recognized, and semantically verified to generate intelligent image and text data; rotating the pen tail triggers the recording function, and the audio is converted into intelligent audio and text data through noise reduction, feature extraction, and speech recognition. The main control chip aligns the timestamps of the image and audio data through a dynamic time warping algorithm to generate intelligent synchronization set data, and then removes redundant text through a cosine similarity algorithm to form an intelligent text set. Finally, Bluetooth Based on the GATT protocol, the module encapsulates multimodal data into binary frames with CRC checksums, and transmits them to the mobile phone APP in real time for parsing into JSON format. Its innovations are: 1. A dual-trigger mechanism (pen tip and finger pressure) to achieve multi-scene image acquisition; 2. A dynamic time warping algorithm to solve the problem of multimodal data synchronization; 3. An efficient data transmission architecture based on the GATT protocol; 4. Multi-dimensional data cleaning and feature enhancement technology to ensure information accuracy and completeness. This acquisition and processing method design breaks through the single-function limitations of traditional pen-type devices and builds a new type of intelligent input terminal that integrates written records, voice transcription, and intelligent synchronization, facilitating daily information collection and organization. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The disclosure of the present invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. In the accompanying drawings, the same reference numerals are used to refer to the same components. Among them: Figure 1 Schematically shows a flow chart of a smart pen data collection and processing method proposed in accordance with one embodiment of the present invention; Figure 2 Schematically shows a detailed process (pen pressure) diagram of step S1 proposed according to one embodiment of the present invention; Figure 3 Schematically shows a detailed process (finger pressure) diagram of step S1 proposed according to one embodiment of the present invention; Figure 4 Schematically shows a detailed flow chart of step S2 proposed according to one embodiment of the present invention; Figure 5 Schematically shows the overall structure of a smart pen proposed according to one embodiment of the present invention; Figure 6 The figure schematically shows the internal structure of a smart pen according to one embodiment of the present invention.
[0017] Numbers in the figure: 1. Smart pen body; 11. Pen tip pressure sensor; 12. OCR camera group; 13. Pen tail microphone; 14. Bluetooth module; 15. Main control chip; 16. Finger pressure sensor. DETAILED DESCRIPTION
[0018] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, a person skilled in the art can propose a variety of interchangeable structural modes and implementation modes. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention and should not be regarded as the entire invention or as a limitation or restriction of the technical solution of the present invention.
[0019] According to one embodiment of the present invention, Figure 1 as well as Figure 5-Figure 6 , a smart pen data acquisition and processing method based on a multimodal sensor, comprising the following steps: S1: By squeezing the pen tip pressure sensor 11 and the finger pressure sensor 16, a corresponding trigger signal is obtained, and based on the trigger signal, the OCR camera group 12 is controlled to turn on, and the content in front of the pen tip is photographed to obtain image data, and the image data is stored; S2: Trigger the recording switch by rotating the pen end, turn on the microphone to record, obtain audio data and store it; S3: The main control chip 15 integrates and processes the stored image data and audio data to obtain and store integrated data; S4: The stored image data, audio data and integrated data are transmitted to the mobile phone APP in real time via the Bluetooth module 14.
[0020] Through a dual trigger mechanism using either a pen tip pressure sensor or a finger pressure sensor, the OCR camera group is controlled to capture and store images of the writing area. To achieve continuous writing signal acquisition, the smart pen design also incorporates a timer to achieve a delay effect. When the pen tip or side is pressed, the camera starts recording and the timer begins. If the signal is interrupted, the timer begins counting. During the set delay, even if the signal disappears, recording continues until the delay ends and there is no new signal. Rotating the pen tail triggers the microphone recording function, synchronously recording audio data. The main control chip intelligently processes the image and audio data. Images are recognized by OCR to generate graphic data, and audio is converted into audio data through noise reduction and speech recognition. A dynamic time warping algorithm is then used to align the timestamps of the graphic and audio data, forming a synchronized data set. Finally, the data is transmitted in real time to a mobile phone app for analysis and processing via a Bluetooth module, achieving the integrated functions of writing, voice transcription, and intelligent synchronization, breaking through the limitations of the single nature of traditional pen-type devices.
[0021] The method is applied to a smart collection pen, which includes a smart pen body 1, a pen tip pressure sensor 11, an OCR camera group 12, a pen tail microphone 13, a Bluetooth module 14, a main control chip 15, an analog-to-digital converter, and a finger pressure sensor 16. The main control chip 15 is provided inside the smart pen body 1, the pen tip pressure sensor 11 is provided at one end of the smart pen body 1, and the finger pressure sensor 16 is provided on the outside of the smart pen body 1. The finger pressure sensor 16 is provided at one end close to the pen tip pressure sensor 11. The OCR camera group 12 is also provided at one end of the smart pen body 1, and the OCR camera group 12 is provided on the outside of the pen tip pressure sensor 11. The pen tail microphone 13 is provided at the other end of the smart pen body 1, the Bluetooth module 14 is provided inside the smart pen body 1, and the main control chip 15 is provided. It is electrically connected to the pen tip pressure sensor 11, the OCR camera group 12, the pen tail microphone 13, the Bluetooth module 14 and the finger pressure sensor 16. A battery is provided inside the smart pen body 1, and a charging interface is provided on the outside of the smart pen body 1 corresponding to the battery. The battery is provided on the side close to the pen tail microphone 13. When using the smart pen, you can write normally like an ordinary pen. When writing, the pen tip pressure sensor 11 is squeezed and the OCR camera group 12 is automatically turned on to take pictures and record the writing and circled content. For some electronic devices that are inconvenient to write on, you can press the finger pressure sensor 16 with your finger to turn on the OCR camera group 12 for recording. In addition, you can rotate the pen tail switch as needed to turn on the pen tail microphone 13 for recording to realize information collection.
[0022] The present invention will be further described below with reference to the embodiments.
[0023] See also Figures 1-4 , S1 includes the following steps: S11: When the user writes normally, the pen tip is squeezed to trigger the pen tip pressure sensor 11. The pen tip pressure sensor 11 collects the pen tip pressure data of the writing process of the smart collection pen in real time to obtain the smart collection pen tip pressure trigger signal; S12: The tip pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the tip pressure sensor 11. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the smart collection pen digital signal. S13: The smart collection pen digital signal is subjected to a filtering algorithm to remove noise to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip 15. The smart collection pen digital control signal is transmitted to the OCR camera group 12 to activate the camera hardware and add an initial image timestamp. S14: The optical image data captured by the OCR camera group 12 is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered, and converted into digital image data through an analog-to-digital converter. Finally, it is compressed and encoded by the main control chip 15 and written to the storage medium, and an end image timestamp mark is added. When the user writes normally, the smart collection pen starts the image collection process through the pen tip pressure sensor. First, the pen tip pressure sensor collects the pen tip pressure data in real time and obtains a pressure trigger signal. This signal is converted into an analog voltage signal. After amplification, filtering, and analog-to-digital conversion, it becomes a digital signal. Then, the digital signal is filtered and denoised. Based on this, the main control chip generates a digital control signal, activates the OCR camera group, and adds an initial image timestamp. Subsequently, the optical image data captured by the OCR camera group is converted into an analog electrical signal through the image sensor. Then, it is amplified, filtered, and converted into digital image data. Finally, the main control chip compresses and encodes it and writes it to the storage medium, while adding an end image timestamp.
[0024] S1 also includes the following steps: S15: When it is inconvenient to squeeze the pen tip when collecting information on the electronic device, the user presses the finger pressure switch to trigger the finger pressure sensor 16. The finger pressure sensor 16 collects finger pressure data of the pressing process of the smart collection pen in real time to obtain a finger pressure trigger signal of the smart collection pen; S16: The finger pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the finger pressure sensor 16. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the smart collection pen digital signal; S17: The smart collection pen digital signal is subjected to a filtering algorithm to remove noise to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip 15. The smart collection pen digital control signal is transmitted to the OCR camera group 12 to activate the camera hardware and add an initial image timestamp. S18: The optical image data captured by the OCR camera group 12 is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered and converted into digital image data through an analog-to-digital converter. Finally, it is compressed and encoded by the main control chip 15 and written into the storage medium, and an end image timestamp mark is added. When it is inconvenient to squeeze the pen tip to collect information on the electronic device, the user presses the finger pressure switch to trigger the finger pressure sensor. The finger pressure sensor collects finger pressure data in real time, obtains a pressure trigger signal, converts it into an analog voltage signal, amplifies and filters it and converts it into a digital signal through analog-to-digital conversion. The digital signal is then filtered and denoised. The main control chip generates a digital control signal based on this, activates the OCR camera group and adds an initial image timestamp. The optical image data captured by the OCR camera group is processed into digital image data, compressed and encoded by the main control chip and stored in the storage medium, and an end image timestamp is added at the same time.
[0025] S2 includes the following steps: S21: The user rotates the recording switch at the end of the pen, and the mechanical action signal generated by the recording switch is converted into a trigger signal for the smart acquisition laptop through the Hall sensor; S22: The trigger signal of the smart collection laptop is shaped and amplified by the signal conditioning circuit and then transmitted to the main control chip 15. After receiving the trigger signal of the smart collection laptop, the main control chip 15 generates a recording control signal of the smart collection pen. The recording control signal of the smart collection pen is transmitted to the microphone 13 at the end of the pen to start the recording function and add an initial recording timestamp mark. S22: The audio sound wave data recorded by the pen tail microphone 13 is converted into an audio analog electrical signal through the electret capacitor. The audio analog electrical signal is converted into digital audio data by the analog-to-digital converter after pre-amplification and filtering, and transmitted to the main control chip 15 through the audio transmission line. It is compressed and encoded by the main control chip 15 and written into the storage medium, and a termination recording timestamp is added. The user rotates the pen tail recording switch, and the mechanical action signal generated is converted into an electrical trigger signal by the Hall sensor. The signal is shaped and amplified by the signal conditioning circuit and then transmitted to the main control chip. After receiving the signal, the main control chip generates a recording control signal. This signal enables the pen tail microphone to turn on the recording function and adds an initial recording timestamp. The audio sound wave data recorded by the pen tail microphone is first converted into an analog electrical signal, and then obtained as digital audio data through pre-amplification, filtering and analog-to-digital conversion. It is then compressed and encoded by the main control chip and stored in the storage medium, and a termination recording timestamp is added.
[0026] S18 includes the following steps: S181: The digital image data is processed by grayscale, noise reduction, and binarization to optimize the image quality and obtain pre-processed image data; S182: Pre-processing the image data to extract text features through an image recognition module to obtain image and text information data; S183: Graphic information data is processed through character correction and semantic verification calculations to obtain intelligent graphic data. A series of processing is performed on digital image data to obtain intelligent graphic data. First, grayscale, noise reduction and binarization operations are performed on the digital image data to optimize the image quality and obtain pre-processed image data. Then, text features are extracted from the pre-processed image data with the help of an image recognition module to obtain graphic information data. Finally, character correction and semantic verification calculations are performed on the graphic information data to remove erroneous or inaccurate content, and finally intelligent graphic data is obtained to provide high-quality graphic information for subsequent use.
[0027] S22 includes the following steps: S221: Converting the digital audio data from the time domain to the frequency domain through Fourier transform to obtain noise audio data; S222: The digital audio data is processed by an adaptive filtering algorithm, and filtering is performed according to the determined noise audio data to remove noise interference and obtain filtered audio data; S223: The filtered audio data is compared with a set energy threshold to detect blank sound segments in the audio, which are then deleted from the filtered audio data to obtain streamlined audio data. The digital audio data is processed to improve quality. First, a Fourier transform is used to convert the audio information from the time domain to the frequency domain, thereby obtaining noisy audio data. An adaptive filtering algorithm is then used to filter the digital audio data based on the obtained noisy audio data to remove noise interference and obtain filtered audio data. Finally, the filtered audio data is compared with the set energy threshold to detect and delete blank sound segments, thereby obtaining streamlined audio data and making the audio purer and more effective.
[0028] S22 further includes the following steps: S224: Simplify the audio data by extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients to obtain audio feature data; S225: The audio feature data is processed and analyzed by matching with the acoustic model and the language model to obtain transcribed text data; S226: The transcribed text data is processed through character correction and semantic verification calculations to obtain intelligent audio-text data, and the simplified audio data is further processed to generate intelligent audio-text data. First, the Mel-frequency cepstral coefficients and linear prediction cepstral coefficients are extracted from the simplified audio data to obtain audio feature data that can reflect the audio characteristics. Then, the audio feature data is matched and analyzed with the acoustic model and language model to convert the audio information into transcribed text data. Finally, character correction and semantic verification calculations are performed on the transcribed text data to correct errors and improve semantics, thereby obtaining accurate and reliable intelligent audio-text data.
[0029] S3 includes the following steps: S31: Integrate the intelligent graphic data and the intelligent audio data to obtain text data set data; S32: The text data set is analyzed and processed using the cosine similarity algorithm and the edit distance algorithm to calculate the similarity of the text information in the data set, find duplicate text content, and obtain redundant text data; S33: Delete redundant text data from the text dataset to obtain intelligent text set data. To obtain the intelligent text set data, first integrate the intelligent image and text data with the intelligent audio and text data to generate the text dataset data. Then, use the cosine similarity algorithm and the edit distance algorithm to perform similarity calculations on the text information in the text dataset, accurately locate the repeated text content, and form redundant text data. Finally, remove the redundant text data from the text dataset, and successfully obtain the deduplicated and more refined intelligent text set data, thereby improving the quality and availability of the data.
[0030] S31 includes the following steps: S311: Collecting the initial image timestamp, the final image timestamp, the initial recording timestamp, and the final recording timestamp in the intelligent image and text data and the intelligent audio and text data; S312: Comparing the intelligent text set data with the intelligent audio and text data and segmenting the time series, segmenting the time series of the intelligent audio and text data and the intelligent text set data; S313: Use the dynamic time warping algorithm to match and align the segmented intelligent audio and text data and the intelligent text set data, adjust the timestamp, and obtain the intelligent synchronization set data. To achieve time synchronization between the intelligent audio and text and the graphic data, first collect the initial and ending images and recording timestamp marks in the intelligent graphic and intelligent audio and text data. Then, compare the intelligent text set data with the intelligent audio and text data, segment the time series of the two, and divide the data reasonably according to the time characteristics. Finally, use the dynamic time warping algorithm to match and align the segmented intelligent audio and text and intelligent text set data, and adjust the timestamp according to the matching situation, so as to obtain the intelligent synchronization set data that can reflect the time correlation of the data.
[0031] S4 includes the following steps: S41: The smart collection pen completes pairing and establishes a connection with the mobile phone APP through the Bluetooth module 14. At the same time, the smart collection device organizes and encapsulates the smart synchronization set data, smart text set data, smart image and text data, and smart audio and text data into binary protocols. By adding a data type identification header, a data length field, and a CRC check code, different types of data are packaged according to a frame structure to prepare for data transmission; S42: The intelligent acquisition device sends the encapsulated data to the mobile phone APP in real time through the Bluetooth module 14. After receiving the data, the mobile phone APP receives the data by subscribing to the characteristic value based on the Bluetooth GATT protocol, calls the corresponding parser according to the predefined frame header identifier, and reorganizes the data into JSON format for processing. When the intelligent acquisition pen and the mobile phone APP transmit data, they first complete the pairing and establish a connection through the Bluetooth module. At the same time, the intelligent acquisition device organizes the intelligent synchronization set, intelligent text set, intelligent image and text, and intelligent audio and text data, encapsulates them according to the binary protocol, adds the data type identification header, data length field and CRC check code, and packages the data in a frame structure. After that, the intelligent acquisition device sends the encapsulated data to the mobile phone APP in real time through the Bluetooth module. The mobile phone APP receives the data by subscribing to the characteristic value based on the Bluetooth GATT protocol, parses it according to the frame header identifier, and then reorganizes it into JSON format for processing.
[0032] In this embodiment, the smart pen is based on multimodal sensor fusion technology to build an innovative information collection terminal that integrates writing records, voice transcription, and intelligent synchronization. Its core architecture includes hardware modules such as the pen tip pressure sensor, OCR camera group, pen tail microphone, and finger pressure sensor, which cooperate with the main control chip to realize the full process intelligence of data collection, processing and transmission.
[0033] In the data collection process, a dual trigger mechanism is adopted: when writing normally, squeezing the pen tip triggers the pressure sensor, generating an analog voltage signal that is converted into a digital control signal, activating the OCR camera group to capture the image and add a timestamp; when operating on an electronic device, pressing the finger pressure sensor triggers the same process to achieve multi-scene adaptability. Audio collection triggers the Hall sensor by rotating the mechanical switch at the end of the pen, converting the motion signal into an electrical signal to control the microphone recording and synchronously record the timestamp.
[0034] The data processing stage includes multi-dimensional intelligent algorithms: after image data is pre-processed by grayscale conversion, noise reduction, and binarization, it is used to generate graphic data through OCR recognition, and character correction and semantic verification are performed; audio data uses Fourier transform to separate noise, combined with adaptive filtering and energy threshold detection technology to remove interference, and extracts features such as Mel-frequency cepstral coefficients. After that, audio and text data is generated through matching of acoustic models and language models. The main control chip uses the cosine similarity algorithm to integrate graphic and audio data, and realizes timestamp alignment through the dynamic time warping algorithm to generate intelligent synchronization set data. The edit distance algorithm is used to remove redundant text to form a refined intelligent text set.
[0035] Data transmission adopts an efficient architecture based on the Bluetooth GATT protocol: the device encapsulates multimodal data into binary frames with CRC checksum, transmits them to the mobile phone app in real time through the characteristic value subscription mechanism, and realizes multi-terminal data synchronization after parsing into JSON format. This design breaks through the single function limitation of traditional writing tools. The innovation lies in: the dual trigger mechanism solves the problem of multi-scenario input, the dynamic time warping algorithm realizes accurate synchronization of multimodal data, the GATT protocol ensures low power and efficient transmission, and the multi-dimensional data cleaning technology improves information accuracy, making it easy to collect and organize information in daily life.
[0036] The technical scope of the present invention is not limited to the contents of the above description. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.
Claims
1. A smart pen data acquisition and processing method based on a multimodal sensor, characterized by: The following steps are involved: S1: By squeezing the pen tip pressure sensor and finger pressure sensor, a corresponding trigger signal is obtained. Based on the trigger signal, the OCR camera group is controlled to start, and the content in front of the pen tip is photographed to obtain image data, and the image data is stored; S2: Trigger the recording switch by rotating the pen end, turn on the microphone to record, obtain audio data and store it; S3: The stored image data and audio data are integrated and processed by the main control chip to obtain integrated data and store it; S4: The stored image data, audio data and integrated data are transmitted to the mobile phone APP in real time through the Bluetooth module.
2. The method for collecting and processing smart pen data based on a multimodal sensor according to claim 1, characterized in that: Said S1 comprises the following steps: S11: When the user writes normally, the pen tip is squeezed to trigger the pen tip pressure sensor, and the pen tip pressure sensor collects the pen tip pressure data of the writing process of the smart collection pen in real time to obtain the smart collection pen tip pressure trigger signal; S12: The tip pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the tip pressure sensor of the pen. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the digital signal of the smart collection pen. S13: The smart collection pen digital signal is filtered to remove noise through a filtering algorithm to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip. The smart collection pen digital control signal is transmitted to the OCR camera group to activate the camera hardware and add an initial image timestamp. S14: The optical image data captured by the OCR camera group is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered, and converted into digital image data through an analog-to-digital converter. The digital image data is finally compressed and encoded by the main control chip and written to the storage medium, and an end image timestamp is added.
3. The method for collecting and processing smart pen data based on a multimodal sensor according to claim 1, wherein: Said S1 further comprises the following steps: S15: When it is inconvenient to squeeze the pen tip when collecting information on the electronic device, the user presses the finger pressure switch to trigger the finger pressure sensor, and the finger pressure sensor collects finger pressure data of the pressing process of the smart collection pen in real time to obtain a finger pressure trigger signal of the smart collection pen; S16: The finger pressure trigger signal of the smart collection pen is converted into an analog voltage signal by the finger pressure sensor. The analog voltage signal is amplified and filtered, and then the analog-to-digital converter obtains the digital signal of the smart collection pen. S17: The smart collection pen digital signal is subjected to a filtering algorithm to remove noise to obtain a filtered smart collection pen digital signal. The filtered smart collection pen digital signal is used to generate a smart collection pen digital control signal through the main control chip. The smart collection pen digital control signal is transmitted to the OCR camera group to activate the camera hardware and add an initial image timestamp. S18: The optical image data captured by the OCR camera group is converted into an analog electrical signal through the image sensor. The analog electrical signal is amplified, filtered, and converted into digital image data through an analog-to-digital converter. Finally, the main control chip compresses and encodes the digital image data and writes it to the storage medium, adding an end image timestamp.
4. The method for collecting and processing smart pen data based on a multimodal sensor according to claim 1, wherein: The S2 comprises the following steps: S21: The user rotates the recording switch at the end of the pen, and the mechanical action signal generated by the recording switch is converted into a trigger signal for the smart acquisition laptop through the Hall sensor; S22: The trigger signal of the smart collection laptop is shaped and amplified by the signal conditioning circuit and then transmitted to the main control chip. After receiving the trigger signal of the smart collection laptop, the main control chip generates a recording control signal of the smart collection pen. The recording control signal of the smart collection pen is transmitted to the microphone at the end of the pen to start the recording function and add an initial recording timestamp mark at the same time; S22: The audio sound wave data recorded by the microphone at the end of the pen is converted into an audio analog electrical signal through the electret capacitor. The audio analog electrical signal is converted into digital audio data by the analog-to-digital converter after pre-amplification and filtering, and transmitted to the main control chip through the audio transmission line. The main control chip compresses and encodes the data and writes it to the storage medium, and adds a timestamp to mark the end of recording.
5. The method for collecting and processing smart pen data based on a multimodal sensor according to claim 3, characterized in that: The S18 comprises the following steps: S181: The digital image data is processed by grayscale, noise reduction, and binarization to optimize the image quality and obtain pre-processed image data; S182: Preprocessing the image data to extract text features through an image recognition module to obtain graphic information data; S183: The graphic information data is processed through character correction and semantic verification calculation to obtain intelligent graphic data.
6. The multimodal sensor-based smart pen data acquisition and processing method according to claim 4, characterized in that: The S22 includes the following steps: S221: Converting the digital audio data from the time domain to the frequency domain through Fourier transform to obtain noise audio data; S222: The digital audio data is processed by an adaptive filtering algorithm, and filtering is performed according to the determined noise audio data to remove noise interference and obtain filtered audio data; S223: The filtered audio data is compared with a set energy threshold to detect blank sound segments in the audio, and the blank sound segments are deleted from the filtered audio data to obtain simplified audio data.
7. The multimodal sensor-based smart pen data acquisition and processing method according to claim 6, characterized in that: The S22 further comprises the following steps: S224: Simplify the audio data by extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients to obtain audio feature data; S225: The audio feature data is processed and analyzed by matching with the acoustic model and the language model to obtain transcribed text data; S226: The transcribed text data is processed through character correction and semantic verification calculations to obtain intelligent audio-text data.
8. The multimodal sensor-based smart pen data acquisition and processing method according to claim 1, characterized in that: The S3 includes the following steps: S31: Integrate the intelligent graphic data and the intelligent audio data to obtain text data set data; S32: The text data set is analyzed and processed using the cosine similarity algorithm and the edit distance algorithm to calculate the similarity of the text information in the data set, find duplicate text content, and obtain redundant text data; S33: Delete redundant text data from the text data set to obtain intelligent text set data.
9. The method for collecting and processing smart pen data based on a multimodal sensor according to claim 8, characterized in that: The S31 includes the following steps: S311: Collecting the initial image timestamp, the final image timestamp, the initial recording timestamp, and the final recording timestamp in the intelligent image and text data and the intelligent audio and text data; S312: Comparing the intelligent text set data with the intelligent audio and text data and segmenting the time series, segmenting the time series of the intelligent audio and text data and the intelligent text set data; S313: Use the dynamic time warping algorithm to match and align the segmented intelligent audio and text data and the intelligent text set data, adjust the timestamps, and obtain the intelligent synchronization set data.
10. The smart pen data acquisition and processing method based on a multimodal sensor according to claim 1, characterized in that: The S4 comprises the following steps: S41: The smart collection pen completes pairing and establishes a connection with the mobile phone app through the Bluetooth module. At the same time, the smart collection device organizes and encapsulates the smart synchronization set data, smart text set data, smart image and text data, and smart audio and text data into binary protocol. By adding a data type identification header, a data length field, and a CRC check code, different types of data are packaged according to the frame structure to prepare for data transmission. S42: The intelligent data collection device sends the encapsulated data to the mobile phone APP in real time through the Bluetooth module. After receiving the data, the mobile phone APP receives the data by subscribing to the characteristic value based on the Bluetooth GATT protocol, calls the corresponding parser according to the predefined frame header identifier, and reorganizes the data into JSON format for processing.