Control method, electronic equipment and handwriting pen equipment

By monitoring user behavior data and automatically switching the recording device mode, the problem of manual operation of existing recording devices is solved, and efficient recording and accurate storage of intelligent recordings are achieved.

CN120669948APending Publication Date: 2025-09-19LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510713639.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In meetings, classrooms and other scenarios, participants or students are unable to fully record the content. Existing recording equipment requires manual opening and closing of recording, which can easily lead to missing important information.

Method used

By monitoring user behavior data, automatically switching the recording device mode, listening to and saving key sound data, and combining body movement data to achieve intelligent recording.

Benefits of technology

It improves the completeness and accuracy of recording content, reduces user manual operations, and improves learning and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669948A_ABST
    Figure CN120669948A_ABST
Patent Text Reader

Abstract

The invention discloses a control method, electronic equipment and handwriting pen equipment, and the method comprises the steps: monitoring user behavior data obtained through target input equipment, the user behavior data comprising sound data and / or limb movement data generated by at least one user object in a space environment collected by the target input equipment; under the condition that the user behavior data meets the target triggering condition, the target input device is controlled to be switched between a first mode and a second mode; wherein the target input device only monitors the sound data in the first mode, and the target sound data in the sound data can be stored in the second mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a control method, an electronic device, and a stylus device. Background Art

[0002] In meetings, classes, interviews, and other scenarios, participants or students often don't have time to record the entire content and can only take shorthand or keyword notes. However, this method results in participants or students being unable to fully grasp the entire content of the meeting or class due to missing notes when reviewing the meeting or reviewing for an exam. Therefore, participants or students may use voice recorders or smart recording software to record the voice content of meetings, classes, interviews, and other scenarios. However, existing voice recorders or smart recording software often rely on manual activation and deactivation of the recording function, which can easily miss important information. Summary of the Invention

[0003] The present application provides a control method, an electronic device, and a stylus device.

[0004] The technical solution of this application is achieved as follows:

[0005] In the first aspect, an embodiment of the present application provides a control method, including: monitoring user behavior data obtained through a target input device, the user behavior data including sound data and / or body movement data generated by at least one user object in the spatial environment collected by the target input device; when the user behavior data meets the target trigger condition, controlling the target input device to switch between a first mode and a second mode; wherein, the target input device only monitors sound data in the first mode and can save target sound data in the sound data in the second mode.

[0006] In the second aspect, an embodiment of the present application provides an electronic device, comprising at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following: monitoring user behavior data obtained through a target input device, the user behavior data including sound data and / or body motion data generated by at least one user object in a spatial environment collected by the target input device; controlling the target input device to switch between a first mode and a second mode when the user behavior data meets a target trigger condition; wherein the target input device only monitors sound data in the first mode and can save target sound data in the sound data in the second mode.

[0007] In a third aspect, an embodiment of the present application provides a stylus device, comprising a device body and an audio monitoring component and a controller arranged in the device body; wherein the device body is capable of collecting body motion data of at least one user object in a spatial environment; the audio monitoring component is used to monitor sound data in the spatial environment; the controller is used to control the stylus device to switch between a first mode and a second mode when user behavior data composed of body motion data and / or sound data meets a target trigger condition; wherein the stylus device only monitors sound data in the first mode, and can save target sound data in the sound data to a memory in the second mode.

[0008] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A schematic diagram of an implementation flow of a control method provided in an embodiment of the present application;

[0010] Figure 2 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0011] Figure 3 A schematic diagram of the structure of a stylus device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0013] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0015] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as generally understood by those skilled in the art in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0016] The present application embodiment provides a control method, such as Figure 1 As shown, it includes step S110 and step S120:

[0017] Step S110: monitoring user behavior data obtained through the target input device, where the user behavior data includes sound data and / or body movement data generated by at least one user object in the spatial environment collected by the target input device;

[0018] Here, the target input device refers to a device with intelligent sensing capabilities that integrates audio collection functions. User behavior data refers to data related to user interaction collected through the target input device, including but not limited to sound data and / or body movement data emitted by the user.

[0019] In some embodiments, the target input device may be a stylus, a touch pen, or a stylus with a recording function, the target input device may also be a handwriting board, a touch pad, or a drawing board with a recording function, the target input device may also be a mouse or a keyboard with a recording function, the target input device may also be a handle or a glove with a recording function, the target input device may also be an image acquisition component with a recording function, and so on.

[0020] In some embodiments, user behavior data may be operation behavior data of a user using a target input device to input, for example, the operation behavior data of input may be writing behavior data of a user using a stylus; user behavior data may also be operation behavior data of a user using a target input device to abandon input, for example, the operation behavior data of abandoning input may be pen-lifting behavior data or pen-putting behavior data of a user pausing writing; user behavior data may also be user image data collected by an image acquisition device, for example, user image data may be gesture operations of a collected user object, user image data may also be voice operations of a collected user object, user image data may also be other body movements of a collected user object, etc., wherein the target input device may be integrated with an image acquisition device, or the image acquisition device may send the collected user behavior data to the target input device; user behavior data may also be sound data collected by a microphone in the target input device when the user speaks, etc.

[0021] In some embodiments, a user object may include a person and / or a device, wherein a user object may include a person with specific characteristics, for example, a person who is speaking, or a pre-designated person; a device may include an intelligent device, for example, a user object may be an audio playback device, or an intelligent device such as a robot.

[0022] In some implementations, when the user object is a person, the sound data refers to language data of the person.

[0023] In some implementations, when the user object is an audio playback device, the sound data may be audio data played by the audio playback device.

[0024] In some embodiments, the body movement data may be body movements with specific characteristics, such as specific gestures, writing movements, drawing movements of a user object, specific interaction movements between a user object and a target input device, and the like.

[0025] It is understood that the present application can trigger the mode switch of the target input device through sound data; or trigger the mode switch of the target input device through body movement data; or trigger the mode switch of the target input device through sound data and body movement data. For example, sound data can be used to trigger a stylus pen, smart keyboard and mouse, smart ring, smart camera, etc. with recording function to switch to recording mode or to a mode of only listening without saving the recording.

[0026] Step S120: When the user behavior data meets the target trigger condition, controlling the target input device to switch between the first mode and the second mode;

[0027] The target input device only monitors the sound data in the first mode, and can save the target sound data in the sound data in the second mode.

[0028] Here, the first mode refers to the target input device only monitoring the surrounding sound data and performing real-time analysis on the sound data without saving the monitored sound data. The second mode refers to the target input device not only monitoring the surrounding sound data, but also performing real-time analysis on the sound data and saving the key sound data it monitors. The target sound data refers to the key sound data in the sound data, for example, the target sound data can be key knowledge points, key summaries, and other language content.

[0029] In some embodiments, the target trigger condition may include but is not limited to at least one of the following: preset keywords; preset specific voiceprint information; preset specific tone information; preset specific gestures; preset specific postures; preset specific body movements.

[0030] In one example, the sound data is matched with preset keywords. If the match is successful, it is determined that the sound data meets the target trigger condition, where the keywords may include but are not limited to any of the following words: focus, attention, summary.

[0031] In one example, voice processing is performed on sound data to obtain voiceprint information corresponding to the sound data; the voiceprint information is matched with preset specific voiceprint information, and if the match is successful, it is determined that the sound data meets the target trigger condition. For example, if a specific person starts speaking, the target input device can be triggered to switch to recording mode and start saving the recording.

[0032] In one example, parameter recognition processing is performed on sound data to obtain intonation information corresponding to the sound data; the intonation information is matched with preset specific intonation information. If the match is successful, it is determined that the sound data meets the target trigger condition. For example, if the speaker suddenly raises the pitch, the target input device can be triggered to switch to recording mode and start saving the recording.

[0033] In one example, image data of a user object is collected through an image acquisition component, and the collected image data is matched with a preset specific gesture or specific posture. If the match is successful, it is determined that the image data meets the target trigger condition. For example, if the speaker switches the PPT to the next page or starts to talk about the target file, the target input device can be triggered to switch to recording mode and start saving the recording.

[0034] In one example, when the user object performs a specific body movement through the target input device, it is determined that the target trigger condition is met, and the target input device is triggered to switch modes. For example, when the user object performs a writing operation or a drawing operation through a stylus or a writing pen, the stylus or writing pen can be triggered to record and save.

[0035] It can be understood that when the target input device is in the first mode, the target input device only monitors the sound data but does not save the collected sound data, and when the target input device is in the second mode, the target input device not only monitors the sound data but also saves the collected sound data. Therefore, the power consumption of the target input device in the first mode is lower than that in the second mode.

[0036] In some embodiments, after monitoring the sound data, the target input device analyzes and processes the sound data using a processing model built into the target input device to generate key sound data in the sound data, namely, target sound data.

[0037] In the embodiments of the present application, on the one hand, by real-time monitoring and analysis of user behavior data, the target input device can be automatically switched between the first mode and the second mode, and the automatic recording operation of the current scene can be realized, so that the user does not need to manually record the target input device, reducing the frequency of manual operation by the user, so that the user can focus more on understanding and recording the current voice content. On the other hand, since the user may miss or mistakenly miss the voice of the current scene when manually operating the target input device to record, the present application automatically controls the target input device to perform the recording operation, which can improve the completeness and accuracy of the recorded content.

[0038] In some embodiments, the control method may further include step S130:

[0039] Step S130: Processing the target sound data and the text data and / or image data formed based on the body movement data stored in the target input device into a target file.

[0040] In some implementations, the target file may be generated by associating target sound data with text data and / or image data.

[0041] In some embodiments, the target file may be obtained by performing text processing on text data and / or image data using target sound data, wherein the text processing may include at least one of the following: optimization processing, completion processing, and regeneration processing.

[0042] In some implementations, the target file may be generated by personalizing the target sound data through text data and / or image data.

[0043] In some embodiments, the target file is also configured with index data that can access the target sound data. In this way, when the user reads the text content in the target file, he can also use the index data to obtain the corresponding target sound data, and with the help of the target sound data, enhance his understanding of the text content.

[0044] In an embodiment of the present application, by integrating the target sound data with the text data and / or image data formed based on the body movement data and processing them into a target file, a seamless combination of intelligent shorthand and voice content is achieved, thereby improving the completeness and accuracy of the content of the target file finally obtained, thereby significantly improving the learning and work efficiency of users in scenarios such as meetings and classrooms.

[0045] In some embodiments, the above step S130, processing the target sound data and the text data and / or image data formed based on the body movement data stored in the target input device into a target file, may further include at least one of the following steps S131 to S134:

[0046] Step S131: inputting target sound data and text data and / or image data into a first processing model for generation processing to obtain a target file;

[0047] Here, the first processing model refers to a comprehensive processing model that integrates speech recognition, natural language processing and multimodal fusion capabilities, and is used to uniformly process speech, text and image data and generate structured content output.

[0048] In some embodiments, the first processing model can be based on a deep learning network, for example, a Transformer architecture, which can simultaneously process multiple types of data inputs and integrate them into a coherent text or document form.

[0049] In some embodiments, when the first processing model receives the target sound data, text data and / or image data, it performs content analysis on the target sound data, text data and / or image data, respectively, and reconstructs the data based on the analyzed content to generate a target file.

[0050] Step S132: supplementing, optimizing, or regenerating the text data and / or image data with reference to the target sound data using the second processing model to obtain a target file;

[0051] Here, the second processing model is an auxiliary processing model based on context perception and content enhancement, which can supplement, optimize or regenerate text data and / or image data based on the existing text data and / or image data through additional information provided by sound data.

[0052] In one example, when a user attends a meeting and records the keywords and / or key sentences of the meeting using a stylus device, the second processing model can automatically complete each keyword and / or key sentence based on the target sound data; the second processing model can also automatically optimize and modify each keyword and / or key sentence based on the target sound data; the second processing model can also automatically correct the recorded keywords and / or key sentences based on the target sound data; the second processing model can also generate text content for the content that the user omitted to record through the target sound data.

[0053] It is understandable that the target sound data can be used to supplement, optimize or regenerate text data and / or image data, which not only makes up for the deficiencies of the original text data and / or image data, but also improves the readability of the content of the text data and / or image data.

[0054] Step S133: performing inference processing on the text data and / or image data using the third processing model to obtain user portrait information of the first user object, and generating and processing the target sound data based on the user portrait information to obtain a target file;

[0055] Here, the third processing mode is a model that focuses on user behavior modeling and personalized content generation. Its core function is to extract the user's behavior patterns, interest preferences and expression style from the input text data and / or image data to construct a digital portrait of the user.

[0056] In some embodiments, the third processing model can be generated based on a machine learning algorithm, such as cluster analysis, deep neural network, etc.

[0057] In one example, when a user records keywords and / or key sentences of a meeting using a stylus device while attending a meeting, the third processing model can infer based on the keywords and / or key sentences currently recorded by the first user object to generate user portrait information of the first user object.

[0058] In one example, when a user attends a meeting and records the keywords and / or key sentences of the meeting through a stylus device, the third processing model can obtain the stored historical user portrait information of the first user object, and update the historical user portrait based on the keywords and / or key sentences currently recorded by the first user object to generate user portrait information of the first user object.

[0059] It can be understood that the target file generated by the third processing model is adapted to the user's note-taking style, which can reduce the difficulty of the user's understanding of the target file, making it easier for the user to master and understand the target file, and meeting the user's personalized needs for the target input device.

[0060] In some embodiments, if the user portrait information indicates that the first user object tends to use a concise and clear language style, the third processing model generates a concise and clear target file based on the content of the target sound data, making it closer to the user's writing habits.

[0061] In some embodiments, since the user's language style may vary in different scenarios, the third processing model can also identify the current scenario in which the user is located; and based on the current scenario, perform inference processing on the text data and / or image data to obtain user portrait information in the current scenario, so that the generated target file is adapted to the current scenario.

[0062] In one example, if the current scenario is a classroom scenario, and if the user portrait information indicates that the first user object tends to use academic textual expressions, the third processing model generates a classroom note file with academic expressions based on the content of the target sound data, making it closer to the user's language style in the classroom scenario.

[0063] In one example, if the current scene is a meeting scene, if the user portrait information represents that the first user object prefers a clear and coherent textual expression of the meeting content, the third processing model generates a clear and coherent meeting minutes document based on the content of the target sound data, making it closer to the user's language style in the meeting scene.

[0064] Step S134: generating and processing the multimedia file displayed by the electronic device using the target sound data stored in the target input device to obtain the target file.

[0065] In some implementations, the target sound data may be used to complete the content of the multimedia file to obtain the target file.

[0066] In one example, when the multimedia file is a PPT file, the target sound data can be used to complement and optimize the content of the PPT file.

[0067] In some implementations, the target sound data can be used to mark key content in a multimedia file to obtain a target file.

[0068] In one example, when the multimedia file is a PPT file, the target sound data can be used to highlight the content in the PPT file.

[0069] In some implementations, the target input device may further associate the target sound data with content in the multimedia file to obtain a target file.

[0070] In one example, when the multimedia file is a PPT file, the sound data in the target sound data can be bound to the content of each page in the corresponding PPT file, so that when the user plays the slide show, the language clips related to the page can be played synchronously, thereby enhancing the user's understanding and memory of the target file.

[0071] In an embodiment of the present application, by processing the target sound data and text data and / or image data into a target file, intelligent fusion of multi-source information is achieved. The target file generated by combining the multi-source information can improve the accuracy and completeness of the content of the target file.

[0072] In some embodiments, the above step S130, processing the target sound data and the text data and / or image data generated based on the body movement data stored in the target input device into a target file, may further include at least one of the following steps S135 and S136:

[0073] Step S135: performing feature extraction processing on the text data and / or image data generated by the handwriting operation of the first user object on the stylus device using the fourth processing model;

[0074] Here, the fourth processing model is an artificial intelligence model specifically used to recognize and parse text data and / or image data. Feature extraction processing refers to the process of extracting key information from text data and / or image data.

[0075] In some embodiments, the fourth processing model is used to perform feature extraction processing on text data and / or image data to generate target feature data, for example, extracting the semantic content or user intention represented by the text data and / or image data, or extracting keyword content or key phrases in the text data, extracting key image elements in the image data, etc. as feature data.

[0076] Step S136: Target associated data is extracted from the target sound data using the extracted target feature data, so as to supplement or optimize the text data and / or image data using the target associated data to obtain a target file.

[0077] In some implementations, the target feature data may be keywords or key image elements extracted from text data and / or image data.

[0078] In some embodiments, the target feature data may also be summary information, outline structure information, mind map information, image description information, image preview data, or image cover data generated based on text data and / or image data.

[0079] In some implementations, the target feature data may also be semantically related features generated based on text data and / or image data.

[0080] It can be understood that the target feature data is generated based on the key information in the text data and / or image data, and then the target associated data corresponding to the target feature data is screened in the target sound data. Through this method, the sound data in the target sound data that is not related to the target feature data is no longer considered, which reduces the data processing amount of the target sound data, improves the processing efficiency of the target sound data, and realizes accurate and fast optimization processing of the text data and / or image data, thereby improving the generation speed of the target file and reducing the time the user waits for the generation of the target file.

[0081] In some embodiments, when the text data includes multiple unrelated text segments recorded by the user, the multiple text segments can be associated with each other through the target feature data to generate a well-organized and logically reasonable target file.

[0082] In some embodiments, when the text data includes multiple interrelated text segments recorded by the user, the multiple interrelated text segments can be verified through the target feature data, and the erroneous information in the text segments can be corrected to generate a correct target file.

[0083] In some implementations, when the text data includes multiple related text keywords for recording, the multiple text keywords can be combined using the target feature data to generate a logically correct text segment.

[0084] In some embodiments, when the text data includes a text outline recorded by the user, the text outline can be supplemented with details using the target feature data to generate a target document with complete content and clear explanations.

[0085] In some embodiments, when the image data includes an image outline recorded by a user, the image outline can be supplemented with details using the target feature data to generate a target file with complete content.

[0086] In some implementations, the target feature data can also be used to verify the image data, correct erroneous image content in the image data, and generate a correct target file.

[0087] In some implementations, the target feature data can also be used to supplement the text description of the image data to generate a target file with complete content and clear explanations.

[0088] In some embodiments, the above step S120, controlling the target input device to switch between the first mode and the second mode, includes at least one of the following steps S121 and S122:

[0089] Step S121: When the user behavior data satisfies a first trigger condition, controlling the target input device to switch from the current first mode to the second mode, wherein satisfying the first trigger condition includes at least one of the following: the first user object uses the target input device to perform a target input operation, the target input device monitors target sound data, and the target input device collects target content data output by the electronic device;

[0090] In some implementations, executing the target input operation may be an operation in which the user performs handwriting input via a target input device. In this scenario, the target input device may be a stylus pen, a handwriting pen, or the like.

[0091] In some implementations, executing the target input operation may be an operation in which the user performs voice input through a target input device.

[0092] In some implementations, executing the target input operation may be an operation in which the user performs a spatial gesture input through a target input device.

[0093] In some implementations, the target sound data may be target content data, such as a trigger of a keyword.

[0094] In one example, when the keywords include words such as key points and summary, for example, when the target input device detects "What I am going to say next is the key points, everyone please listen carefully", it determines that the keyword has been triggered, and then controls the target input device to switch to the second mode.

[0095] In some implementations, the target sound data may be target voiceprint data.

[0096] In one example, the target voiceprint data may be that of a specific speaking object, and the target input device is provided with the voiceprint data of the specific speaking object, so that the target input device can match the voiceprint data of the monitored sound data with the voiceprint data of the specific speaking object. If the match is successful, the target input device is controlled to switch to the second mode.

[0097] In some implementations, the target sound data may be target tone data.

[0098] In one example, a target tone threshold is set in the target input device. This allows the target input device to compare the tone data of the monitored sound data with the target tone threshold. If the tone data of the sound data exceeds the target tone threshold, the target input device is controlled to switch to the second mode. It is understandable that in this scenario, the target tone threshold is used to determine whether the first trigger condition is met because when a speaker speaks to a key point, they may increase their voice to attract the audience's attention.

[0099] In some embodiments, the target content data output by the electronic device may be the target multimedia file currently displayed by the electronic device, such as a PPT file about the theme of the current conference, such as product discussions and plans for exhibition at the Mobile World Congress or the Consumer Electronics Show.

[0100] In some implementations, the electronic device outputting the target content data may be the electronic device currently sharing specific information.

[0101] It can be understood that this application sets multiple first trigger conditions so that the target input device can flexibly match the corresponding first trigger conditions in different scenarios according to the user behavior in the current scenario. In this way, the recording saving function can be automatically turned on without manual intervention by the user, thereby avoiding missing important content, reducing the user's operating burden, and improving overall efficiency.

[0102] Step S122: When the user behavior data meets the second trigger condition, control the target input device to switch from the current second mode to the first mode, wherein meeting the second trigger condition includes at least one of the target input device completing the execution of the target input operation, the target input device monitoring the first sound data, and the electronic device completing the output of the target content data.

[0103] In some implementations, completing the target input operation may be when the user stops performing the handwriting input operation through the target input device.

[0104] In one example, when it is detected that the user is performing a pen-down operation or a pen-down operation, it is determined that the user stops performing the handwriting input operation through the target input device.

[0105] In some implementations, completing the target input operation may be when the user stops performing the voice input operation through the target input device.

[0106] In some implementations, completing the target input operation may be when the user stops performing the spatial gesture input operation through the target input device.

[0107] In some implementations, the first sound data may be other voiceprint data that is unrelated to the target voiceprint data.

[0108] In one example, the first voiceprint data may be voiceprint data of another object. It is understandable that in this scenario, the specific speaking object has stopped speaking and another object is speaking.

[0109] In one example, the first voiceprint data may be voiceprint data of the current location. It is understandable that in this scenario, the specific speaker has stopped speaking, and no one is speaking at the current location.

[0110] In some implementations, the first sound data may be a first tone data that is significantly lower than the target tone data.

[0111] In one example, a first tone threshold is set in the target input device. The target input device can compare the tone data of the monitored sound data with the first tone threshold. If the tone data of the sound data is less than the first tone threshold, the target input device is controlled to switch to the first mode. It is understandable that in this scenario, the first tone threshold is used to determine whether the second trigger condition is met because the speaker may lower the volume when speaking non-essential content.

[0112] In one example, a first tone threshold is set in the target input device, so that the target input device can compare the tone data of the monitored sound data with the first tone threshold, and when the tone data of the sound data is less than the first tone threshold and maintains the first duration, the target input device is controlled to switch to the first mode.

[0113] In some implementations, the first sound data may be first content data, such as a trigger of a first keyword.

[0114] In one example, when the first keyword includes words such as non-emphasis and off-topic, for example, when the target input device detects "What I am going to say next is the non-emphasis content of this section", it determines that the first keyword has been triggered, and then controls the target input device to switch to the first mode.

[0115] In some implementations, the electronic device completing the output of the target content data may mean that the current electronic device ends displaying a multimedia file, such as a PPT file.

[0116] In some implementations, the electronic device completing outputting the target content data may be the electronic device ending sharing of specific information.

[0117] It can be understood that the present application sets multiple second trigger conditions so that the target input device can flexibly match the corresponding second trigger conditions in different scenarios according to the user behavior in the current scenario. In this way, the recording saving function can be automatically turned off without manual intervention by the user, thereby reducing the storage data writing of the target input device, saving the storage space and power consumption of the target input device, and effectively extending the battery life of the target input device.

[0118] In some embodiments, the above step S121, controlling the target input device to switch from the current first mode to the second mode, includes the following steps S1211:

[0119] Step S1211: when the first user object is monitored to perform a handwriting operation using the stylus device or the stylus device monitors target sound data, controlling the stylus device to switch from only monitoring the sound data in the spatial environment to executing a saving operation on the monitored target sound data;

[0120] and / or,

[0121] The stylus device can establish a data transmission channel with the electronic device in the first mode to transmit the monitored sound data to the electronic device;

[0122] In the second mode, the data transmission channel between the stylus device and the electronic device is disconnected so that the target sound data is saved in its own storage space.

[0123] Here, the stylus device is an intelligent pen-type input device integrated with a microphone and storage functions, which not only supports handwriting input but also can collect surrounding sound information in real time.

[0124] In some embodiments, the first user object uses a stylus device to perform a handwriting operation, which can be using the stylus device to perform a writing operation in an input area of ​​the electronic device (such as a touchpad or touch screen), or using the stylus device to perform a writing operation in three-dimensional space, such as spatial writing within a spatial range, or writing or drawing on a plane within the space.

[0125] In one example, a first user performs a writing operation on a plane in a physical space by using a stylus device.

[0126] In some implementations, the stylus device is controlled to switch from only monitoring the sound data in the spatial environment to performing a saving operation on the monitored target sound data, that is, the stylus device is controlled to switch from the current first mode to the second mode.

[0127] In some embodiments, when the stylus device is in the first mode, the stylus device can only monitor the sound of the current scene. In this scenario, a data transmission channel is established between the stylus device and the electronic device, and the monitored sound data is sent to the electronic device, so that the electronic device saves and / or processes the received sound data.

[0128] It should be noted that when the stylus device is in the first mode, establishing a data transmission channel between the stylus device and the electronic device is an optional solution. That is to say, when controlling the stylus device to switch from only monitoring the sound data in the spatial environment to executing the saving operation of the monitored target sound data, if there is a data transmission channel between the stylus device and the electronic device, it is also necessary to disconnect the data transmission channel between the stylus device and the electronic device to save the target sound data to the storage space of the stylus device itself; or, when controlling the stylus device to switch from only monitoring the sound data in the spatial environment to executing the saving operation of the monitored target sound data, if there is no data transmission channel between the stylus device and the electronic device, there is no need to disconnect the data transmission channel between the stylus device and the electronic device.

[0129] In some embodiments, when the stylus device is in the second mode, the data transmission channel between the stylus device and the electronic device is disconnected, and the target sound data is saved in the stylus device's own memory. In this way, when the target sound data is used to generate a target file, the stylus device can read the target sound data stored in itself without the need to transmit the target sound data, thereby improving the efficiency of target file generation.

[0130] In some embodiments, when the stylus device is switched from the second mode back to the first mode, the stylus device can also send the saved target sound data to the electronic device, and back up the target sound data through the electronic device. In this way, when the stylus device is abnormal, the target sound data stored in the electronic device can be obtained, thereby realizing the processing of the target sound data.

[0131] In some embodiments, when the stylus device is switched from the second mode back to the first mode, the stylus device can also send the saved target file to the electronic device, and back up the target file through the electronic device. In this way, when the stylus device malfunctions, the target file stored in the electronic device can be obtained, thereby realizing the processing of the target file.

[0132] In an embodiment of the present application, when the target input device meets the requirements of switching to the first mode, the stylus device is automatically triggered to record. This reduces the frequency of manual operations of the first user object, effectively solves the problem of missed or incorrect recordings caused by manual operations, and improves the completeness and accuracy of the recorded content.

[0133] In some embodiments, the above step S120, controlling the target input device to switch between the first mode and the second mode, includes the following step S123:

[0134] Step S123: When the user behavior data satisfies the first trigger condition, controlling the target input device to process the monitored target sound data using the target artificial intelligence service; and

[0135] When the user behavior data satisfies the second trigger condition, the obtained processing result is added to the text data and / or image data formed by the target input operation.

[0136] Here, artificial intelligence services refer to intelligent systems integrated into target input devices that can analyze and process target sound data based on speech recognition, semantic understanding, and natural language processing technologies.

[0137] It can be understood that when the user behavior data meets the first trigger condition, the target input device enters the first mode. In this scenario, the target input device stores the monitored target sound data and processes the target sound data through the target artificial intelligence service.

[0138] In some implementations, the processing result is obtained based on a summary of the target sound data.

[0139] In one example, the target artificial intelligence service may include an artificial intelligence summary function, so that when the target input device enters the first mode, the monitored target sound data is summarized in real time through the artificial intelligence summary service, and when the target input device enters the second mode, the summarized content is converted into text and / or image to obtain the processing result.

[0140] In some implementations, the processing result is obtained based on text and / or image conversion of the target sound data.

[0141] In one example, when the target input device enters the first mode, the target sound data is converted into text and / or image in real time through the target artificial intelligence service, and the processing result is obtained when the target input device enters the second mode.

[0142] In some embodiments, when the target input device enters the second mode, it indicates that the current target sound data has been recorded. In this scenario, the processing result can be text data and / or image data obtained by text and / or image conversion based on the target sound data. In this way, the text data or image data obtained after the conversion is added to the text data and / or image data formed by the user performing the target input operation, thereby improving the integrity and correctness of the content of the target file finally formed.

[0143] In other embodiments, when the target input device enters the second mode, it indicates that the current target sound data has been recorded. In this scenario, the processing result can also be text data obtained after summarizing the target sound data (such as a speech outline or summary or emphasized part obtained by sorting out the speaker's speech content) and / or image data (including image parameters, such as image style, image color parameters, image specification parameters, etc. obtained by analyzing the speaker's speech content). In this way, the text data or image data obtained after the summary can reduce the content added to the text data and / or image data formed by the user performing the target input operation, or supplement or modify the text data and / or image data formed by the target input operation, thereby improving the conciseness and completeness of the content of the target file finally formed.

[0144] In some embodiments, based on the user portrait information, the processing result is converted into a language style represented by the user portrait information; and the converted processing result is added to the text data and / or image data, so that the final target file can be adapted to the user's writing habits.

[0145] In some embodiments, when the obtained processing results are added to text data and / or image data, a link can also be generated for the target sound data corresponding to the processing results and inserted into the corresponding text data and / or image data, so that the user can obtain the target sound data by clicking on the link.

[0146] In the embodiments of this application, a target AI service is introduced to intelligently process the target sound data. When specific trigger conditions are met, the processing results are embedded into the text data and / or image data generated by the user's target input operation. This not only effectively reduces the user's manual operation burden, but also improves the integrity and accuracy of the text data and / or image data, improves the user's learning and work efficiency, and significantly improves the user's learning and work experience.

[0147] In some embodiments, during the execution of step S123, any one of the following steps S124 and S125 may also be included:

[0148] Step S124: presenting a process of adding the processing result to the text data and / or image data on the target interactive interface of the electronic device, wherein the display parameters of the processing result and the text data and / or image data are different;

[0149] In some embodiments, the display parameters may include a main body color, so that the processing results and text data and / or image data can be displayed in different colors to allow users to distinguish them.

[0150] In some embodiments, the display parameters may include identification parameters, so that the processing results can be highlighted through the identification parameters, or the text data and / or image data can be highlighted through the identification parameters to enable the user to distinguish them.

[0151] In one example, the identification parameter may include highlighting the processing result, so that the processing result can be distinguished from text data and / or image data through the identification parameter.

[0152] In one example, the identification parameter may include setting an underline below the processing result, so that the processing result can be distinguished from text data and / or image data through the identification parameter.

[0153] In one example, the identification parameter may include bolding the processing result, so that the processing result can be distinguished from text data and / or image data through the identification parameter.

[0154] It can be understood that the display parameters of the processing results are different from those of the text data and / or image data. In this way, the user can distinguish between the processing results and the text data and / or image data through the target interactive interface. This differentiated display method can help users more intuitively distinguish between automatically generated content and input content, improve the user's reading efficiency and understanding of the final file content, and at the same time reduce the possibility of users misunderstanding and confusion about the final file content, thereby enhancing the user's usage experience.

[0155] Step S125: providing an editable option control for the processing result added to the text data and / or image data, so as to update the display parameters and / or display content of the target processing result in response to the configuration operation of the editable option control by the first user object.

[0156] Here, the editable option control refers to a button or menu item that allows the user to select and modify the target processing result on the interface.

[0157] In some implementations, the target processing result is generated after the processing result is added to the text data and / or image data.

[0158] It is understood that users can use the editable option controls to adjust the display parameters and / or display content of the target processing results according to their own needs, so that the adjusted target processing results are more in line with the user's personal habits and / or expression methods, thereby improving the readability of the target processing results. In addition, the editable option space also enhances the system's personalized adaptation capabilities.

[0159] The embodiment of the present application provides an electronic device 200, such as Figure 2As shown, it includes at least one processor 1 and at least one processing model 2 that can run on the processor 1, and the processing model can be called by the target application to perform at least one of the following: monitoring user behavior data obtained through the target input device, the user behavior data including sound data and / or body movement data generated by at least one user object in the spatial environment collected by the target input device; when the user behavior data meets the target trigger condition, controlling the target input device to switch between the first mode and the second mode; wherein, the target input device only monitors the sound data in the first mode and can save the target sound data in the sound data in the second mode.

[0160] In some embodiments, the target input device and the electronic device may be communicatively connected, and the target input device transmits the monitored user behavior data to the electronic device via a communication connection link.

[0161] In one example, the target input device can communicate with the electronic device via a wired connection. In another example, the target input device can communicate with the electronic device via a Bluetooth connection.

[0162] In some implementations, the electronic device is provided with application software adapted to the target input device, and interaction with the target input device is achieved through the application software in the electronic device.

[0163] In some embodiments, the processing model can be called by the target application to execute: processing the target sound data stored by the target input device and the text data and / or image data formed based on the body motion data into a target file.

[0164] In some embodiments, the target input device sends the saved target sound data to the electronic device, and the target input device sends the text data and / or image data formed by the body motion data to the electronic device, and the processing model of the electronic device processes the target sound data and the text data and / or image data formed by the body motion data into a target file.

[0165] In some embodiments, the processing model can be called by the target application to perform at least one of the following: when the user behavior data meets the first trigger condition, the target input device is controlled to switch from the current first mode to the second mode, wherein satisfying the first trigger condition includes at least one of the first user object using the target input device to perform the target input operation, the target input device listening to the target sound data, and the target input device collecting the target content data output by the electronic device; when the user behavior data meets the second trigger condition, the target input device is controlled to switch from the current second mode to the first mode, wherein satisfying the second trigger condition includes at least one of the target input device completing the execution of the target input operation, the target input device listening to the first sound data, and the electronic device completing the output of the target content data.

[0166] In some embodiments, when the target input device is a stylus device, the processing model can be called by the target application to perform at least one of the following: when the first user object is monitored to perform a handwriting operation using the stylus device or the stylus device listens to target sound data, the stylus device is controlled to switch from only listening to the sound data in the spatial environment to performing a saving operation on the listened target sound data; and / or, the stylus device can establish a data transmission channel with the electronic device in a first mode to transmit the listened sound data to the electronic device; in a second mode, the data transmission channel between the stylus device and the electronic device is disconnected to save the target sound data in its own storage space.

[0167] In some embodiments, the processing model can be called by the target application to perform at least one of the following: when the user behavior data meets the first trigger condition, control the target input device to use the target artificial intelligence service to process the monitored target sound data; and when the user behavior data meets the second trigger condition, add the obtained processing results to the text data and / or image data formed by the target input operation.

[0168] In some embodiments, the electronic device also includes a display component, which is used to present the processing process of adding the processing result to the text data and / or image data through the target interactive interface; the processing model can be called by the target application to execute: presenting the processing process of adding the processing result to the text data and / or image data on the target interactive interface of the electronic device, wherein the processing result is different from the display parameters of the text data and / or image data; providing an editable option control for the processing result added to the text data and / or image data, so as to respond to the configuration operation of the editable option control by the first user object, and execute the update of the display parameters and / or display content of the target processing result.

[0169] In some embodiments, the processing model includes one or more of the following: at least one of a first processing model, a second processing model, and a third processing model; the processing model can be called by the target application to perform at least one of the following: inputting the target sound data and text data and / or image data into the first processing model for generation processing to obtain a target file; using the second processing model to supplement, optimize, or regenerate the text data and / or image data with reference to the target sound data to obtain a target file; using the third processing model to infer the text data and / or image data to obtain user portrait information of the first user object, and generating the target sound data based on the user portrait to obtain a target file; using the target sound data stored in the target input device to generate and process the multimedia file displayed on the electronic device to obtain a target file.

[0170] In some embodiments, the processing model includes a fourth processing model, which can be called by the target application to perform at least one of the following: using the fourth processing model to perform feature extraction processing on the text data and / or image data generated by the handwriting operation of the first user object on the stylus device; using the extracted target feature data to extract target association data from the target sound data, and using the target association data to supplement or optimize the text data and / or image data to obtain the target file.

[0171] The present application embodiment provides a stylus device, such as Figure 3 As shown, it includes a device body 3 and a memory 4, an audio monitoring component 5 and a controller 6 arranged in the device body 3; wherein the device body 3 can collect body motion data of at least one user object in the spatial environment; the audio monitoring component 5 is used to monitor the sound data in the spatial environment; the controller 6 is used to control the stylus device to switch between the first mode and the second mode when the user behavior data composed of the body motion data and / or the sound data meets the target trigger condition; wherein the stylus device only monitors the sound data in the first mode, and can save the target sound data in the sound data to the memory in the second mode.

[0172] In some implementations, the audio monitoring component may be a microphone, an audio capture card, or other device capable of monitoring and capturing audio.

[0173] In some implementations, the body movement data of the user holding the stylus device is acquired through a sensor and / or an image acquisition device provided on the stylus device.

[0174] In one example, the sensor may be a pressure sensor, a six-degree-of-freedom sensor, or the like.

[0175] In one example, the image acquisition device may be a video camera, a still camera, or the like.

[0176] In some embodiments, the image acquisition device provided on the stylus device is used to obtain the body movement data of the user subject who does not hold the stylus device.

[0177] In some embodiments, the controller is configured to process the target sound data and the text data and / or image data formed based on the body motion data stored in the stylus device into a target file.

[0178] In some embodiments, the controller is used for at least one of the following: controlling the stylus device to switch from the current first mode to the second mode when the user behavior data satisfies a first trigger condition, wherein satisfying the first trigger condition includes at least one of the first user object using the stylus device to perform a target input operation, the stylus device listening to target sound data, and the stylus device collecting target content data output by the electronic device; controlling the stylus device to switch from the current second mode to the first mode when the user behavior data satisfies a second trigger condition, wherein satisfying the second trigger condition includes at least one of the stylus device completing the execution of the target input operation, the stylus device listening to the first sound data, and the electronic device completing the output of the target content data.

[0179] In some implementations, the stylus device is communicatively connected to the electronic device, so that the stylus device can collect target content data output by the electronic device.

[0180] In some embodiments, the stylus device is connected to the electronic device for communication, and the stylus device can send the display status of the current target content data to the stylus device through the communication connection link.

[0181] In some embodiments, the controller is used to: when monitoring a first user object using a stylus device to perform a handwriting operation or the stylus device listens to target sound data, control the stylus device to switch from only listening to the sound data in the spatial environment to performing a saving operation on the listened target sound data; and / or, the stylus device can establish a data transmission channel with the electronic device in a first mode to transmit the listened sound data to the electronic device; in a second mode, the data transmission channel between the stylus device and the electronic device is disconnected to save the target sound data to its own storage space.

[0182] In some embodiments, the controller is used to: when the user behavior data meets the first trigger condition, control the stylus device to use the target artificial intelligence service to process the monitored target sound data; and when the user behavior data meets the second trigger condition, add the obtained processing results to the text data and / or image data formed by the target input operation.

[0183] In some embodiments, the controller is further configured to at least one of: send text data and / or image data to the electronic device, and send a processing result to the electronic device.

[0184] In some embodiments, when an electronic device receives text data and / or image data, as well as a processing result, it presents a processing process of adding the processing result to the text data and / or image data on a target interactive interface of the electronic device, wherein the processing result is different from the display parameters of the text data and / or image data; and provides an editable option control for the processing result added to the text data and / or image data, so as to update the display parameters and / or display content of the target processing result in response to a configuration operation of the editable option control by a first user object.

[0185] In some embodiments, the controller is used for at least one of the following: inputting target sound data and text data and / or image data into a first processing model for generation processing to obtain a target file; using a second processing model to supplement, optimize, or regenerate text data and / or image data with reference to the target sound data to obtain a target file; using a third processing model to infer text data and / or image data to obtain user portrait information of a first user object, and generating the target sound data based on the user portrait to obtain a target file; using the target sound data stored in a target input device to generate and process multimedia files displayed on an electronic device to obtain a target file.

[0186] In some embodiments, the controller is used to: use a fourth processing model to perform feature extraction processing on text data and / or image data generated by the handwriting operation of the first user object on the stylus device; use the extracted target feature data to extract target association data from the target sound data, and use the target association data to supplement or optimize the text data and / or image data to obtain a target file.

[0187] The following describes the application of the embodiments of the present application in actual scenarios.

[0188] In meetings, classes, interviews and other scenarios, participants or students often do not have time to record the entire content and can only use smart shorthand or keyword recording. As a result, when reviewing the meeting or reviewing for an exam, participants or students are often unable to fully grasp the complete content of the meeting or class due to the lack of notes.

[0189] Currently, the content of meetings, classes, and interviews can be recorded using voice recorders or smart recording software. However, most existing voice recorders or smart recording software rely on manually turning the recording function on and off to record the audio of the entire meeting, class, or interview process. In addition, smart recording software can convert the recorded content into text through voice recognition, but it has problems such as low accuracy, slow recognition speed, sensitivity to environmental noise, and stiff and difficult-to-understand content.

[0190] Therefore, the voice recorders or smart recording software currently available on the market have the following shortcomings: 1) Cumbersome manual operation: Users need to manually control the opening and closing of the recording while listening to the lecture, which can easily lead to missing important content. 2) Redundant information: Recording the entire meeting or course audio results in a large amount of redundant information, which increases the difficulty of subsequent sorting and review. 3) Inaccurate speech recognition: Existing speech recognition technology has low recognition accuracy when processing speech in specific scenarios, such as multi-person communication in meetings or lectures, and environmental noise in closed spaces. 4) Impact on user work and learning efficiency: Users need to record and understand at the same time in meetings or classes. Traditional recording equipment and speech recognition software have failed to effectively improve work and learning efficiency.

[0191] In order to overcome the above problems, an embodiment of the present application provides an AI stylus (i.e., the above-mentioned target input device), which aims to intelligently identify and record important speech parts by integrating AI technology, and combine the user's shorthand content to generate complete, accurate and easy-to-understand records, thereby improving work and learning efficiency.

[0192] First, let me introduce the core functions of this AI stylus.

[0193] 1) Power saving mode and environmental sound collection: After the user starts the AI ​​stylus, first, the AI ​​stylus will enter power saving mode and continuously collect the surrounding environmental sounds; then, the AI ​​stylus will form a contextual environment with the collected voice and input it into the built-in AI system; finally, the AI ​​system will determine whether the subsequent language content is important by analyzing the voice content. For example, when it detects iconic language such as "What I am going to say next is the key point, everyone should listen carefully", the AI ​​stylus will automatically start recording, and judge the time to end the recording based on the voice content, and re-enter power saving mode (that is, the first mode mentioned above) until the next important speech triggers the recording function. At the same time, because the AI ​​stylus can continuously form context through voice, it can better optimize the recording, remove noise, and repair unclear sounds.

[0194] 2) The AI ​​summary and key points are triggered when writing: During the entire recording process, the AI ​​stylus synchronously monitors the changes in the sound waves corresponding to the current ambient sound. When the user starts writing with the AI ​​stylus, the AI ​​summary function is automatically triggered. When the user pauses writing, the AI ​​system will intelligently locate the rising point of the sound wave as the starting point of recording, and continue recording until the end of the sound wave. It will then add key information to the shorthand content according to the user's personal note style, making the user's shorthand content complete and accurate, and attaching a real-time recording of this content to the final shorthand content (ie, the target processing result). The meeting or study notes formed in this way are in the user's own language style, which is easier for users to understand and quickly master.

[0195] 3) Recordings are combined with screen content to improve summaries: The AI ​​system continuously monitors PowerPoint page turning and automatically matches recordings with each PowerPoint page on the timeline. When the user switches to a new page, the system automatically supplements and improves the notes by combining the current PowerPoint content with the corresponding recording, ensuring the generated notes are more accurate and complete.

[0196] Next, the core functions of the AI ​​stylus are explained with reference to its specific implementation principles.

[0197] 1) Sound-based contextual tagging

[0198] Sound wave starting point detection: The AI ​​system monitors sound wave changes (i.e., tone data) through full recording. When a rising point of a sound wave is detected, it marks the rising point as the context starting point.

[0199] During implementation, the AI ​​system uses a root mean square (RMS) algorithm to detect changes in sound waves across several consecutive sampling points, identifying significant increases in volume. The system also filters out background noise using a preset threshold (i.e., the target tone threshold), effectively avoiding recording invalid sound signals.

[0200] Sound wave end point detection: When the sound wave drops to a relatively low threshold (i.e., the first tone threshold) and maintains for a certain period of time, the system determines that the sound wave has ended and then marks the end point as the end point of the recording.

[0201] It should be noted that the AI ​​system uses a preset threshold to distinguish valid speech from noise. When the sound wave intensity falls below this threshold, it is automatically identified as noise. The AI ​​system also analyzes the completeness of the recorded speech content starting from the first valid sound wave. If the content is found to be incomplete, the AI ​​system searches for sound wave segments in the noise preceding the first valid sound wave that can complete the speech content and re-identifies them as valid sound waves.

[0202] 2) Automatically fill in notes with recording content

[0203] Interactive interface design: After the recording is completed, the user can see the voice summary gradually filled in the written text content in the interface. At the same time, the system can also provide manual editing options (that is, the above-mentioned editable option control).

[0204] It should be noted that the voice-filled note content displayed in the interface can be displayed in different colors or special marks so that users can distinguish between their own writing and the content generated by the AI ​​system.

[0205] Underlying Technology: The AI ​​system uses speech recognition technology (for example, the Azure Speech API) to convert the recording into text. Based on the context of the user's written content, it automatically analyzes the parts of the recording that are directly relevant to the user's writing and gradually inserts the converted recording into the user's writing area. Note that this feature can be dynamically updated and displayed using a RichTextBox control or FlowDocument.

[0206] 3) The logic and principle of continuing recording when writing is paused

[0207] The principle of recording continuing during a writing pause is that when the user stops writing (i.e., the target input operation mentioned above), it can be assumed that the user's thoughts have been basically expressed. Therefore, after the user pauses writing, the AI ​​system's AI summary function can be triggered to summarize and output the recorded content and the user's written content. The recording function continues to run during the writing pause, so that the user's subsequent thoughts are not missed. In addition, during the recording pause, the AI ​​system will also convert the recorded content into structured notes.

[0208] It should be noted that triggering the AI ​​summary function means that the AI ​​system will extract content related to the user's written notes from the time when the first sound wave rises before the user puts pen to paper to the time when the first sound wave falls after the user stops writing, and improve and supplement the notes.

[0209] Logical support: During the writing process, users may not be able to quickly write down all the details. After triggering the recording, it can capture additional information when the user is "thinking" or "dictating", and provide an AI-enhanced summary when the user stops writing.

[0210] 4) Achieve adaptive user language style

[0211] ① Implementation that adapts to user style:

[0212] Interactive Interface Design: Users can choose whether to enable adaptive recording to fill in their written notes. In the settings, users can select different style options, such as formal, casual, or professional, so that the final filled note is generated in the user's selected style. In addition, users can edit and correct automatically generated notes. The AI ​​model will simultaneously learn from the user's modification preferences to optimize the note generation effect in the future.

[0213] Underlying technology implementation: Based on the user's long-term note-taking habits, the AI ​​system uses natural language processing technology and machine learning models (such as GPT-3 / 4) to train the user's common phrases, writing style and expression methods, infer the user's language habits (that is, the above-mentioned user portrait information), and insert content of similar style based on the context to adapt to the user's note-taking style.

[0214] It should be noted that the style adaptation function can be continuously strengthened through continuous learning, gradually improving the AI ​​system's ability to understand users' personalized needs.

[0215] ②AI personalized learning:

[0216] The user can manually modify each note generated, and the system will adjust the style of the AI-generated notes based on the user's modification records.

[0217] During the text generation process, AI systems can prioritize and adapt to users' commonly used sentence structures and vocabulary choices by adopting language style transfer technology based on Markov Chain or Transformer models. This technology enables AI to generate text that is more in line with users' personal style and preferences.

[0218] Finally, the use process of the above-mentioned AI stylus is described, which may include steps S10 to S15:

[0219] Step S10: Voice input collection.

[0220] Here, when the AI ​​stylus switches to power saving mode, the AI ​​stylus will only retain the data transmission function. In this mode, the AI ​​stylus can capture the surrounding voice information through its built-in microphone and transmit it to the AI ​​system for processing.

[0221] Step S11: speech recognition and semantic understanding.

[0222] In daily use, the AI ​​system will learn to identify language patterns and the user's note-taking style that trigger key information (i.e., the target keywords mentioned above). For example, the model can recognize iconic language such as "What I am going to say next is the key point, everyone should listen carefully." Once this voice content is recognized, the AI ​​system will conduct in-depth semantic understanding and extract the key information.

[0223] Step S12: triggering key language recognition.

[0224] The AI ​​system determines whether the semantic analysis results contain trigger keywords or semantic information, and decides whether to trigger the recording function (i.e., the first trigger condition mentioned above).

[0225] Step S13: Intelligent recording control.

[0226] When the AI ​​system identifies an important speech or the user begins taking notes, the recording function (the second mode mentioned above) is automatically activated. The AI ​​model analyzes the contextual voice information and automatically stops recording at the end of the speech segment that the user deems important, ensuring that the recorded content matches the key points of the meeting or class. In addition, the AI ​​continuously receives contextual information and optimizes the quality of the recorded content to improve the clarity and relevance of the recording.

[0227] Step S14: Complete the association between the notes and the recording.

[0228] Here, the AI ​​system intelligently completes the note content based on the recording content and the user's shorthand style (i.e. the user portrait information mentioned above).

[0229] In addition, the AI ​​model will also associate the recording content with the generated shorthand notes (that is, the above-mentioned target processing results), and there will be a corresponding recording player next to each shorthand note that can be minimized and expanded.

[0230] Step S15: recording and storing.

[0231] Here, the AI ​​system will also generate a link at the end of the note that can jump to the location where the recording is stored, so that the user can obtain the recording related to the note (that is, the target sound data mentioned above) through the link.

[0232] Based on the above embodiments, the following beneficial effects can be achieved:

[0233] 1) Automated recording function: Compared with traditional manual recording devices, the AI ​​stylus can automatically recognize key voice or automatically trigger recording when the user starts writing, eliminating the need for manual operation by the user, reducing tedious operations and the risk of missing important content.

[0234] 2) Combining intelligent completion with shorthand: When the user pauses writing, the AI ​​stylus can simulate the user's note-taking style by combining the recorded content to generate complete shorthand content. The combination of recording and shorthand makes the final record more detailed and accurate, making it easier for users to understand when reviewing.

[0235] 3) Improved Work and Learning Efficiency: Combining real-time recording and intelligent shorthand, users can focus on understanding and recording key points of meeting or class content without having to worry about operating the recording device or waiting for voice recognition transcription, significantly improving their work and learning efficiency. Furthermore, because the AI ​​stylus can simulate the user's shorthand and language style, it generates records that suit their habits and understanding, making review more natural and efficient.

[0236] 4) Improved recording quality: Avoid omissions or misrecordings caused by manual operation, ensuring the completeness and accuracy of the recording content. At the same time, the AI ​​stylus, with its complete context, can optimize the recording content, making it clearer and more accurate.

[0237] 5) Save power: The AI ​​stylus has limited power. Using AI smart recording can extend the standby time of the recorder.

[0238] It should be noted that, in the embodiment of the present application, if the above-mentioned information processing device is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0239] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0240] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes some or all of the steps for implementing the above method.

[0241] The present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a processor, some or all of the steps in the above method are implemented. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).

[0242] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, computer program, and computer program product embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0243] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0244] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0245] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0246] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0247] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0248] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0249] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words be embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0250] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. A control method, comprising: Monitoring user behavior data obtained through a target input device, wherein the user behavior data includes sound data and / or body movement data generated by at least one user object in a spatial environment collected by the target input device; When the user behavior data satisfies a target trigger condition, controlling the target input device to switch between a first mode and a second mode; The target input device only monitors the sound data in the first mode, and can save target sound data in the sound data in the second mode.

2. The method according to claim 1, further comprising: The target sound data stored in the target input device and the text data and / or image data formed based on the body movement data are processed into a target file.

3. The method according to claim 1 or 2, wherein: Controlling the target input device to switch between the first mode and the second mode includes at least one of the following: When the user behavior data satisfies a first trigger condition, controlling the target input device to switch from the current first mode to the second mode, wherein satisfying the first trigger condition includes at least one of the following: the first user object uses the target input device to perform a target input operation, the target input device monitors target sound data, and the target input device collects target content data output by the electronic device; When the user behavior data satisfies the second trigger condition, the target input device is controlled to switch from the current second mode to the first mode, wherein satisfying the second trigger condition includes at least one of the target input device completing the execution of the target input operation, the target input device monitoring the first sound data, and the electronic device completing the output of the target content data.

4. The method according to claim 3, wherein: Controlling the target input device to switch from the current first mode to the second mode includes: When monitoring that the first user object uses the stylus device to perform a handwriting operation or the stylus device monitors target sound data, controlling the stylus device to switch from only monitoring the sound data in the spatial environment to executing a saving operation on the monitored target sound data; and / or, The stylus device is capable of establishing a data transmission channel with the electronic device in the first mode to transmit the monitored sound data to the electronic device; In the second mode, the data transmission channel between the stylus device and the electronic device is disconnected, so that the target sound data is saved in its own storage space.

5. The method according to claim 3, further comprising: When the user behavior data satisfies a first trigger condition, controlling the target input device to process the monitored target sound data using the target artificial intelligence service; as well as, In a case where the user behavior data satisfies a second trigger condition, the obtained processing result is added to the text data and / or image data formed by the target input operation.

6. The method according to claim 5, further comprising at least one of the following: The target interactive interface of the electronic device presents a processing process of adding the processing result to the text data and / or image data, wherein: The processing result is different from the display parameters of the text data and / or image data; An editable option control is provided for the processing result added to the text data and / or image data, so as to update the display parameters and / or display content of the target processing result in response to the configuration operation of the editable option control by the first user object.

7. The method according to claim 2, wherein processing the target sound data stored by the target input device and the text data and / or image data formed based on the body movement data into a target file comprises at least one of the following: Inputting the target sound data and the text data and / or image data into a first processing model for generation processing to obtain the target file; Using a second processing model to supplement, optimize, or regenerate the text data and / or the image data with reference to the target sound data to obtain the target file; Performing inference processing on the text data and / or image data using a third processing model to obtain user portrait information of a first user object, and generating processing on the target sound data based on the user portrait to obtain the target file; The target sound data stored in the target input device is used to generate and process the multimedia file displayed by the electronic device to obtain the target file.

8. The method according to claim 2, wherein processing the target sound data stored by the target input device and the text data and / or image data formed based on the body movement data into a target file comprises: performing feature extraction processing on text data and / or image data generated by the handwriting operation of the first user object on the stylus device using a fourth processing model; Target associated data is extracted from the target sound data using the extracted target feature data, so as to supplement or optimize the text data and / or image data using the target associated data to obtain the target file.

9. An electronic device comprising at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following: Monitoring user behavior data obtained through a target input device, wherein the user behavior data includes sound data and / or body movement data generated by at least one user object in a spatial environment collected by the target input device; When the user behavior data satisfies a target trigger condition, controlling the target input device to switch between a first mode and a second mode; in, The target input device only monitors the sound data in the first mode, and can save target sound data among the sound data in the second mode.

10. A stylus device, comprising a device body and a memory, an audio monitoring component and a controller arranged in the device body; wherein, The device body is capable of collecting body movement data of at least one user object in the spatial environment; The audio monitoring component is used to monitor the sound data in the spatial environment; The controller is used to control the stylus device to switch between the first mode and the second mode when the user behavior data composed of the body motion data and / or the sound data meets the target trigger condition; wherein, the stylus device only monitors the sound data in the first mode and can save target sound data in the sound data to the memory in the second mode.