Recording apparatus, recording system, recording device-based multimodal control management method, electronic device, and medium
By incorporating a microphone component and employing a multimodal control and management method within the recording device, the limitations of microphone design and the singular nature of recording content management in recording equipment are resolved, thereby improving recording quality and user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HANGZHOU DEXIO ROBOTS CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-07-30
AI Technical Summary
The microphone design of existing recording equipment is limited by the internal space structure of the device, resulting in poor recording quality. Furthermore, the control and management of recorded content is limited, leading to a poor user experience.
A recording device is provided, including a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit. The microphone assembly is located inside the housing and is capable of collecting audio from different directions. It is connected to a target terminal through the communication unit and supports multimodal control and management methods to achieve diversified processing of the recorded content.
It improves recording quality and audio capture quality, provides portable recording devices, enhances user experience through multimodal control and management methods, and enables diversified processing of recorded content.
Smart Images

Figure CN2026073468_30072026_PF_FP_ABST
Abstract
Description
Recording apparatus, recording system, multimodal control and management method based on recording equipment, electronic equipment and media
[0001] This application claims priority to Chinese Patent Application No. CN202520175206.1, filed on January 26, 2025, entitled “Recording Device and Recording System”, and Chinese Patent Application No. CN202510125192.7, filed on January 26, 2025, entitled “Multimodal Control Management Method, Electronic Device and Medium Based on Recording Device”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments of this application relate to the fields of recording devices and computers, specifically to recording apparatus, recording systems, multimodal control and management methods based on recording devices, electronic devices, and media. Background Technology
[0003] With the continuous development of artificial intelligence and speech recognition technologies, speech processing technology is being used more and more widely in various electronic products, especially in the field of recording equipment. Currently, most recording devices utilize the built-in recording functions of smart devices (such as mobile phones and tablets), which rely on their built-in microphones. However, when recording content using this method, the following technical problems often arise: the design of the microphones built into most smart devices is often affected by the internal spatial structure of the device, resulting in poor recording quality.
[0004] With the rapid development of information technology, recording equipment is increasingly widely used in various fields, from professional film and television production and news interviews to daily teaching records, meeting minutes, and personal life recordings. Recording equipment plays an indispensable role in these scenarios. At the same time, people are placing increasingly higher demands on the functionality, performance, and ease of use of recording equipment. Currently, most recording content control and management processes involve simply acquiring and storing the content. However, when using this method, the following technical problems often arise: current recording control and management methods are too simplistic and fail to meet diverse needs, such as organizing and marking recorded content, resulting in a poor user experience.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the concept of this application, and therefore may contain information that does not form prior art known to those skilled in the art. Technical issues
[0006] This application provides a recording device and recording system that can collect audio from different directions outside, thereby improving the recording effect and audio acquisition quality.
[0007] This application also provides a multimodal control and management method, electronic device, and medium based on a recording device, which can perform diversified processing on the recorded content and improve the user experience. Technical solutions
[0008] The summary section of this application is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0009] Some embodiments of this application propose recording apparatus and recording systems to solve one or more of the technical problems mentioned in the background section above.
[0010] In a first aspect, some embodiments of this application provide a recording device, wherein the recording device includes a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit, wherein the power management unit and the microphone assembly are disposed within the housing, and the microphone assembly is configured to collect audio from different directions outside; the main control unit and the communication unit are disposed within the housing, and the recording device is configured to communicate with the target terminal via the communication unit.
[0011] In some embodiments of this application, the structure of the recording device includes one of the following: card type, wearable type.
[0012] In some embodiments of this application, the recording device has a card-type structure, and the thickness of the card-type recording device is no more than 1 cm.
[0013] In some embodiments of this application, the size of the card-type recording device covers the set size of the target terminal.
[0014] In some embodiments of this application, the recording device is provided with a magnet, which is disposed on a bonding surface on the outer shell for bonding with the target terminal. The magnet is configured to allow the outer shell to be bonded to the target terminal by magnetic attraction.
[0015] In some embodiments of this application, the recording device further includes a data charging connection cable.
[0016] In some embodiments of this application, the data charging cable includes an external interface and is positioned around the bottom of the housing.
[0017] In some embodiments of this application, a groove is provided at the bottom of the housing, and the data charging connection cable is arranged around the groove.
[0018] In some embodiments of this application, the bottom of the housing is provided with the data charging cable outlet, and the data charging cable can be pulled out of the outlet.
[0019] In some embodiments of this application, at least one button is provided on the aforementioned housing.
[0020] In some embodiments of this application, a touchpad is provided on the aforementioned housing.
[0021] In some embodiments of this application, the touchpad is disposed on the opposite side of the recording device and the target terminal.
[0022] In some embodiments of this application, the communication unit includes at least one of the following: a Bluetooth module, a Wi-Fi module, a 5G module, a 4G module, and an NFC module.
[0023] In some embodiments of this application, the recording device further includes an indicator light, which is disposed in a hole opened in the housing and is communicatively connected to the main control unit.
[0024] In some embodiments of this application, the at least one button includes a power button and a mode switching button, and the touchpad is used to control the setting operation of the recording device.
[0025] In some embodiments of this application, the recording device is further provided with a shooting unit, which is configured to acquire external images.
[0026] In some embodiments of this application, the above-mentioned shooting unit is disposed on the top of the above-mentioned recording device.
[0027] In some embodiments of this application, the recording device further includes a vibration unit, which is communicatively connected to the main control unit; the main control unit is configured to control the vibration unit to perform vibration operations.
[0028] In some embodiments of this application, the bonding surface on the outer shell for bonding with the target terminal is provided with a micro adhesive pad.
[0029] In some embodiments of this application, the microphone assembly includes microphones distributed on at least two sides of the housing.
[0030] In some embodiments of this application, the recording device includes a signal processing unit disposed in the housing, and the signal processing unit includes a processing chip.
[0031] In some embodiments of this application, the power management unit is configured to be connected to the data charging cable, and the external interface of the data charging cable is configured to be connected to a power supply device to supply power and / or charge the power management unit. The power supply device includes at least one of the following: the target terminal and the charging device.
[0032] In some embodiments of this application, the recording device communicates by being attached to the target terminal.
[0033] Secondly, some embodiments of this application provide a recording system, which includes a recording device as described in any of the ways of the first aspect above and a target terminal connected to the recording device, the target terminal being used to display a recording request.
[0034] Thirdly, some embodiments of this application provide a multimodal control and management method based on a recording device. The method includes: in response to detecting a start recording operation through a recording device connected via communication, acquiring first recording content; setting the first recording content through the recording device; and switching different recording modes in response to detecting a preset operation through the recording device.
[0035] In some embodiments of this application, the method further includes: displaying a recording pop-up window in response to determining that a preset recording start condition is met, wherein the recording pop-up window displays a screen recording authorization control; performing screen recording in response to detecting a confirmation operation applied to the screen recording authorization control to obtain screen recording content; and setting the screen recording content.
[0036] In some embodiments of this application, the method further includes: acquiring audio through the audio acquisition unit included in the recording device; preprocessing the audio acquired by the audio acquisition unit to obtain processed recorded audio.
[0037] In some embodiments of this application, the method further includes: generating recording summary information based on the first recording content described above.
[0038] In some embodiments of this application, the above-mentioned setting operation of the first recorded content through the recording device includes: in response to the marking operation detected by the recording device, marking the first recorded content to obtain the first recorded content with marking.
[0039] In some embodiments of this application, the above-mentioned response to detecting a preset operation by the recording device to switch to different recording modes includes: in response to the recording device detecting a mark-end operation, switching to an unmarked recording mode.
[0040] In some embodiments of this application, the above-mentioned marking process for the first recorded content to obtain marked first recorded content includes: parsing the preset target information in the first recorded content to obtain a parsing result; and marking the first recorded content according to the parsing result.
[0041] In some embodiments of this application, the above-mentioned marking process of the first recorded content based on the above-mentioned parsing result includes: in response to determining that the above-mentioned parsing result satisfies the marking triggering condition corresponding to the preset target, marking the preset target in the first recorded content; in response to determining that the obtained parsing result satisfies the marking stop condition corresponding to the preset target, stopping the marking process of the preset target in the first recorded content.
[0042] In some embodiments of this application, the method further includes: generating recording summary information based on the first recorded content marked with tags.
[0043] In some embodiments of this application, the above-mentioned marking process for the first recorded content based on the above-mentioned parsing results includes: performing hierarchical processing on the above-mentioned parsing results to obtain hierarchical parsing results; and marking the first recorded content based on the above-mentioned hierarchical parsing results.
[0044] In some embodiments of this application, the method further includes: generating second recording content based on the first recording content and the obtained screen recording content.
[0045] In some embodiments of this application, the method further includes: storing the second recorded content to the recording device and / or cloud server.
[0046] In some embodiments of this application, generating second recording content based on the first recording content, the screen recording content, or the processed recorded audio includes: performing a fusion process on the first recording content and the screen recording content to obtain the second recording content, wherein the first recording content includes audio of the environment in which the recording device is located, and the second recording content includes the screen recording content and audio corresponding to the environment.
[0047] In some embodiments of this application, the method further includes: in response to meeting the recording start condition of the target meeting, displaying a preset meeting recording pop-up window corresponding to the target meeting, wherein the meeting recording pop-up window displays a screen recording authorization control; in response to detecting a selection operation performed on the displayed screen recording authorization control, screen recording of the target meeting is performed to obtain meeting recording content; and the meeting recording content is determined as screen recording content.
[0048] In some embodiments of this application, the method further includes: in response to detecting that there is an opened setting software content in the screen recording content, displaying a recording pop-up window corresponding to the setting software content, wherein the recording pop-up window displays a screen recording authorization control; in response to detecting a selection operation performed on the displayed screen recording authorization control, screen recording is performed on the setting software content to obtain the recording content of the corresponding setting software as the screen recording content.
[0049] In some embodiments of this application, the method further includes: parsing the recorded content of the corresponding setting software to obtain a content parsing result corresponding to the content of the setting software; and performing a preset operation on the content of the setting software through the recording device based on the content parsing result.
[0050] In some embodiments of this application, the above-mentioned preset operation on the content of the setting software through the recording device based on the content parsing result includes: in response to the detection of a reply operation through the setting software and the setting software being an email type, generating email reply content based on the content parsing result; and inputting the email reply content through the recording device.
[0051] In some embodiments of this application, the above-mentioned preset operation on the content of the setting software through the recording device based on the above-mentioned content parsing result includes: in response to the detection of a reply operation through the above-mentioned setting software, and the above-mentioned setting software being a chat type, generating message reply content based on the above-mentioned content parsing result; and inputting the message reply content through the above-mentioned recording device.
[0052] In some embodiments of this application, the above-mentioned preprocessing of the audio obtained by the audio acquisition unit to obtain processed recorded audio includes: performing noise reduction processing on the audio obtained by the audio acquisition unit to obtain noise-reduced recorded audio; and performing volume adjustment processing on the noise-reduced recorded audio in response to the noise-reduced recorded audio meeting a preset volume difference condition to obtain processed recorded audio.
[0053] In some embodiments of this application, the method further includes: displaying a recording pop-up window in response to detecting proximity information of the recording device, wherein the recording pop-up window displays at least one recording settings control.
[0054] Fourthly, some embodiments of this application provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0055] Fifthly, some embodiments of this application provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above. Beneficial effects
[0056] Some embodiments of this application provide a recording device, offering a more portable recording device with higher recording quality. Specifically, the reason for the low portability and poor recording quality of most recording devices is that most current voice recorders adopt traditional cylindrical or columnar designs. While these designs are fully functional, they often bring inconveniences. For example, most voice recorders cannot be easily carried with common smart devices (such as mobile phones and tablets), limiting their portability. Furthermore, the microphones built into most smart devices are often limited by the internal space structure, resulting in poor recording quality. Based on this, some embodiments of this application provide a recording device including a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit. The power management unit and the microphone assembly are disposed within the housing. The microphone assembly is configured to collect audio from different directions. The main control unit and the communication unit are also disposed within the housing. The recording device is configured to communicate with a target terminal via the communication unit. The aforementioned recording device is an external device, and the microphone assembly is not affected by the internal structural design of the smart device. Furthermore, the microphone assembly is configured to capture audio from different directions, thus improving the recording quality. Therefore, a recording device with improved audio acquisition quality is provided, enhancing the user experience.
[0057] The above embodiments of this application have the following beneficial effects: The multimodal control and management method based on a recording device provided by this application can offer a method for diversified processing of recorded content based on an external recording device. Specifically, the reason why most current methods for managing recorded content are limited to recording and generating content is that users cannot perform further operations on the recorded content, such as marking and organizing. Therefore, some embodiments of this application's multimodal recording method based on a recording device first acquires first recorded content in response to the recording device detecting a start recording operation via a communication connection. This yields preliminary recorded content. Then, the first recorded content is set using the recording device. This allows the user to perform preset operations on the recorded content. Finally, in response to the recording device detecting a preset operation, different recording modes are switched. This process can be achieved by connecting the external recording device to a smart device, and then processing and manipulating the recorded content through the interaction between the smart device and the external recording device. This provides a method for diversified processing of recorded content, improving the user experience. Attached Figure Description
[0058] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0059] Figure 1 is a schematic diagram of the structure of a recording device according to some embodiments of this application;
[0060] Figure 2 is a structural schematic diagram of the recording device of some embodiments of this application in the retracted state of the data charging connection cable;
[0061] Figure 3 is a schematic diagram of the structure of a recording device with a shooting unit according to some embodiments of this application;
[0062] Figure 4 is a schematic diagram of the structure of a recording device installed on a target terminal according to some embodiments of this application;
[0063] Figure 5 is a front view and a rear view of a recording apparatus according to some embodiments of this application;
[0064] Figure 6 is a schematic diagram of another embodiment of the recording device of this application;
[0065] Figure 7 is a structural schematic diagram of the groove cover in the closed state of another embodiment of the recording device of this application.
[0066] Figure 8 is a flowchart of some embodiments of the multimodal control and management method based on a recording device according to this application;
[0067] Figure 9 is a flowchart of some other embodiments of the multimodal control management method based on a recording device according to this application;
[0068] Figure 10 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of this application.
[0069] Implementation methods of this application
[0070] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0071] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0072] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0073] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0074] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0075] The present application will now be described in detail with reference to the accompanying drawings and embodiments.
[0076] Figure 1 is a schematic diagram of the recording device according to some embodiments of this application. Figure 1 includes a housing 1, a data charging connection cable mounting slot 11, a data charging connection cable 2, a button 3, and an indicator light 4.
[0077] Figure 2 is a structural schematic diagram of the recording device of some embodiments of this application in the retracted state of the data charging connection cable. Figure 2 includes a USB connector 21.
[0078] Figure 3 is a schematic diagram of a recording device with a shooting unit according to some embodiments of this application. Figure 3 includes a USB connector 21 and a shooting unit 6.
[0079] Figure 4 is a schematic diagram of the recording device of some embodiments of this application installed on the target terminal. Figure 4 includes the target terminal 5.
[0080] Figure 5 is a front view and a rear view of a recording device according to some embodiments of this application. Figure 5 includes a miniature adhesive pad 13, a magnet 14, and a button 3.
[0081] In some embodiments, the recording device may include a main control unit, a power management unit, a housing 1, a microphone assembly, and a communication unit. The main control unit may be a digital signal processor (DSP) or an application-specific integrated circuit (ASIC), without specific limitation. The main control unit can be used to connect and manage various hardware modules of the recording device, such as sensors, storage devices, and interface circuits. The power management unit may be a rechargeable lithium-ion battery, without specific limitation. The power management unit can be used to supply power to the recording device. The material of the housing 1 may include, but is not limited to, aluminum alloy, ABS engineering plastic, or carbon fiber composite material, without specific limitation. For example, the material of the housing 1 may be a composite structure of aluminum alloy and anti-oxidation plastic, with an outer anti-slip frosted layer, and an overall thickness of 2.5mm to 4mm. The structure of the housing 1 may include, but is not limited to, a plate-like or ring-like shape, as long as it can enclose the internal components without interfering with their functionality, without specific limitation. The housing 1 can be used to protect the components inside the recording device to a certain extent. The microphone assembly may include at least one microphone. For example, each microphone can have a diameter of 2mm and a length of approximately 5mm, and can be manufactured using MEMS (Micro-Electro-Mechanical Systems) technology, featuring high sensitivity and low power consumption. The microphones can be omnidirectional; for example, each omnidirectional microphone can be 3mm in size, and can be arranged in two rows and two columns within the aforementioned housing to form an array for capturing ambient sound in 360°, without specific limitations. The microphone assembly can be used to collect external sounds. The communication unit can include, but is not limited to, Bluetooth, Wi-Fi, 5G, 4G modules, and NFC modules. The communication unit can be used to transmit the audio signals collected by the recording device to the target terminal 5 in real time.
[0082] In some embodiments, the power management unit and the microphone assembly may be housed within the housing 1 to prevent damage to a certain extent. The microphone assembly may be configured to capture audio from different directions to improve the quality of the captured audio.
[0083] In some embodiments, the outer casing 1 may be configured to fit snugly against the target terminal 5, making the recording device more portable. The fitting method may include, but is not limited to, adhesive or magnetic adsorption, etc. The target terminal 5 may include, but is not limited to, a mobile phone, tablet, or computer. The mobile phone may include mobile terminals and foldable phones, which are not specifically limited here.
[0084] In some embodiments, the main control unit and the communication unit may be housed within the housing 1 to prevent damage to a certain extent. The recording device is configured to communicate with the target terminal 5 via the communication unit.
[0085] In some embodiments of this application, the structure of the recording device may include one of the following: card-type or wearable. The wearable type may include, but is not limited to, watch-type and bracelet-type. Compatibility with different structures can make the recording device more versatile.
[0086] In some embodiments of this application, the recording device can be card-shaped. The thickness of the card-shaped recording device can be no more than 1 cm. This thickness of no more than 1 cm allows the recording device to have a thinner and lighter form, making it easier for users to carry.
[0087] In some embodiments of this application, the size of the card-type recording device can cover the set size of the target terminal 5, so as to more flexibly accommodate different target terminals 5. The set size of the target terminal 5 can be the size of the power supply on the back of a mobile phone.
[0088] In some embodiments of this application, as shown in Figures 4-5, the recording device may be provided with a magnet 14. The magnet 14 may be disposed on a bonding surface on the outer casing 1 for attaching with the target terminal 5. The magnet 14 may be configured to magnetically attach the outer casing 1 to the target terminal 5. For example, the magnet may be a magnetic ring, and the charger may be attracted to the magnetic ring. For instance, as shown in Figure 4, the recording device can be attached to the back of a mobile phone via the magnet 14 for easy carrying by the user.
[0089] In some embodiments of this application, as shown in FIG1, the recording device may further include a data charging cable 2. The data charging cable may be a USB cable. The data charging cable 2 may also include a USB connector 21. The data charging cable 2 may be a Type-C cable. The functions of the data charging cable 2 and the USB connector 21 may include, but are not limited to, charging the power management unit and establishing a communication connection with the target terminal 5, without specific limitations. It should be noted that the connection method between the data charging cable 2 and the recording device in FIG1 is only one example; the actual situation can be adjusted as needed. For example, the data charging cable 2 may also be an external cable.
[0090] In some embodiments of this application, as shown in Figures 1-2, the aforementioned data charging cable 2 may include an external interface to enable the recording device to be wiredly connected to the target terminal 5. The data charging cable is positioned around the bottom of the housing. The USB connector 21 may include the external interface.
[0091] In some embodiments of this application, as shown in FIG1, the outer casing 1 may further be provided with a data charging cable mounting slot 11. The data charging cable may be arranged around the mounting slot. The data charging cable mounting slot 11 can be used to store the data charging cable 2 and the USB connector 21, making the recording device more convenient to carry and protecting the USB connector 21 and the data charging cable 2 to a certain extent.
[0092] In some embodiments of this application, the bottom of the housing 1 may be provided with the data charging cable outlet. The data charging cable 2 may be designed to be pulled out of the outlet to increase the flexibility of the data charging cable 2.
[0093] In some embodiments of this application, as shown in Figures 1-2, at least one button 3 may be provided on the housing 1. For example, as shown in Figure 1, two buttons 3 may be provided on the housing 1. The buttons 3 can be used to assist users in performing specified functions, including but not limited to: on / off and pause, etc., without specific limitation. It should be noted that the at least one button is not limited to the physical buttons shown in Figures 1-6, but may also be a touch button or other medium that can implement corresponding commands, without specific limitation here.
[0094] In some embodiments of this application, a touchpad may be provided on the housing 1. The touchpad can set and display relevant functional components of the recording device. These functional components may include, but are not limited to, power buttons and mode switches, etc. It should be noted that the touchpad is not shown in the figures.
[0095] In some embodiments of this application, the touchpad may be positioned opposite the recording device and the target terminal for easy viewing by the user.
[0096] In some embodiments of this application, the communication unit may include at least one of the following: a Bluetooth module, a Wi-Fi module, a 5G module, a 4G module, and an NFC module. It only needs to be able to achieve a communication connection with the target terminal 5, and is not specifically limited. The NFC module of the recording device is used to communicate with the target terminal, establishing a connection through contact, and performing operations according to the operation buttons or touchpad on the recording device. When an NFC communication connection with the target terminal is detected, the recording device is turned on, and a recording pop-up or other operation pop-up will appear on the target terminal. For example, operations such as marking, replying, and summarizing the recording content can also be performed. The recording pop-up is displayed on the target terminal's interface, and remains in the background after the user performs the corresponding recording or canceling operation. To terminate the program, a corresponding operation needs to be performed on the recording device.
[0097] In some embodiments of this application, as shown in FIG1, the recording device may further include an indicator light 4. The indicator light 4 may be an LED, and no specific limitation is made herein. The indicator light 4 may be disposed in a hole opened in the housing 1. The indicator light 4 may be communicatively connected to the main control unit. The indicator light 4 may be used to display the recording status by emitting different colors of light; for example, green light may indicate that recording is in progress, red light may indicate that the device is in standby mode, and blue light may indicate that the device is connected to a mobile phone.
[0098] In some embodiments of this application, the at least one button 3 may include a power button, a mode switching button, and a touchpad for controlling the setting operations of the recording device. The setting operations may include, but are not limited to, turning on, switching, marking, and providing feedback. Both the power button and the mode switching button can be executed via set operation commands. The power button can be used to turn the recording device on and off. For example, the power button can be set to turn on the device with a single press, or to turn on the device with a long press, and to turn off the device when the same operation is performed again. The mode switching button can be used to put the recording device into and out of a marking state. For example, when there is content that needs attention during recording, the device can enter marking mode by double-clicking or long-pressing the mode switching button, and then exit marking mode by long-pressing or double-clicking.
[0099] In some embodiments of this application, as shown in FIG3, the recording device may further be equipped with a shooting unit. The shooting unit may be a miniature camera or a photosensitive element, etc., and is not specifically limited. The shooting unit is configured to acquire external images so that the recording device can perform video recording. It should be noted that the shooting unit is not shown in any figure other than FIG3.
[0100] In some embodiments of this application, as shown in FIG3, the above-mentioned shooting unit can be disposed on the top of the above-mentioned recording device to facilitate the user to acquire images.
[0101] In some embodiments of this application, the recording device may further include a vibration unit. The vibration unit may be a linear resonant actuator (LRA), without specific limitation. The vibration unit may be disposed within the housing 1. The vibration unit is communicatively connected to the main control unit. The main control unit is configured to control the vibration unit to perform vibration operations. For example, when the recording device is started or the recording state changes, feedback is provided through vibration to enhance the user experience.
[0102] In some embodiments of this application, as shown in FIG5, the bonding surface on the outer casing 1 for bonding with the target terminal 5 may be provided with a micro adhesive pad 13. The micro adhesive pad 13 may include, but is not limited to, acrylic adhesive pads or silicone adhesive pads. The micro adhesive pad 13 can be used to ensure that the recording device is more securely bonded to one side of the target terminal 5.
[0103] In some embodiments of this application, the microphones included in the microphone assembly can be distributed on at least both sides of the housing 1 to ensure sound source acquisition from more directions and to a certain extent guarantee the quality of audio acquisition. For example, the microphone assembly can be arranged with two unidirectional microphone units located on the left and right sides of the recording device, forming a simple microphone assembly through a fixed distance arrangement. Alternatively, four omnidirectional microphone units can be used to form a square array layout, distributed at the four corners of the recording device, to capture sound more evenly. It should be noted that the microphone assembly can be located inside the housing 1, but is not shown in the figures.
[0104] In some embodiments of this application, the recording device may include a signal processing unit. The signal processing unit may be housed within the housing 1 to prevent damage to a certain extent. The signal processing unit includes a processing chip. This processing chip may be from the TMS320C55x series, and is not specifically limited thereto. The functions of the signal processing unit may include, but are not limited to, optimizing audio signals, converting and encoding audio formats, etc. It should be noted that the signal processing unit is not shown in the figures.
[0105] In some embodiments of this application, the structure of the outer casing 1 described above may include one of the following: card-type, watch-type, or bracelet-type. It should be noted that Figures 1-7 are merely examples of a card-type structure.
[0106] In some embodiments of this application, the power management unit can be configured to connect to the data charging cable 2. The external interface of the data charging cable 2 can be configured to connect to a power supply device to supply power and / or charge the power management unit. The power supply device includes at least one of the following: the target terminal 5, and a charging device. The charging device may include, but is not limited to, a power bank, a power adapter, etc. For example, the data charging cable 2 can supply power to the power management unit by connecting its external interface to the USB port of a mobile phone or power bank.
[0107] In some embodiments of this application, the recording device communicates by being attached to the target terminal. For example, if the target terminal 5 is a mobile phone or tablet, it can be directly attached and carried; if it is a laptop or the like, it can be directly connected via the data charging cable 2.
[0108] Some embodiments of this application provide a recording device, offering a more portable recording device with higher recording quality. Specifically, the reason for the low portability and poor recording quality of most recording devices is that most current voice recorders adopt traditional cylindrical or columnar designs. While these designs are fully functional, they often bring inconveniences. For example, most voice recorders cannot be easily carried with common smart devices (such as mobile phones and tablets), limiting their portability. Furthermore, the microphones built into most smart devices are often limited by the internal space structure, resulting in poor recording quality. Based on this, some embodiments of this application provide a recording device including a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit. The power management unit and the microphone assembly are disposed within the housing. The microphone assembly is configured to collect audio from different directions. The main control unit and the communication unit are also disposed within the housing. The recording device is configured to communicate with a target terminal via the communication unit. The aforementioned recording device is an external device, and the microphone assembly is not affected by the internal structural design of the smart device. Furthermore, the microphone assembly is configured to capture audio from different directions, thus improving the recording quality. Therefore, a recording device with improved audio acquisition quality is provided, enhancing the user experience.
[0109] Figure 6 is a schematic diagram of another embodiment of the recording device of this application. Figure 5 includes a built-in camera recess 15, a snap-fit recess 151, a first movable hinge 152, an outer bracket 153, a second movable hinge 154, an inner bracket 155, a third movable hinge 156, and a miniature camera 157.
[0110] Figure 7 is a structural schematic diagram of the groove cover in the closed state of another embodiment of the recording device of this application. Figure 6 includes the groove cover 16.
[0111] In some embodiments, as shown in Figures 6-7, the housing may further include a miniature camera 157, a movable camera bracket, a built-in camera recess 15, and a recess cover 16. The miniature camera 157 may be a CMOS miniature camera, without specific limitation. The miniature camera 157 may be wireless or wired. In the wired case, the connecting cable can be disposed within the movable camera bracket to communicate with the recording device. The built-in camera recess 15 may have a locking recess 151 on its periphery. The recess cover 16 may be configured to engage with the built-in camera recess 15 via the locking recess 151, thereby protecting the internal miniature camera 157 and the movable camera bracket to a certain extent. The movable camera bracket may include an outer bracket 153 and an inner bracket 155. The area of the outer bracket 153 is larger than the area of the inner bracket 155. One end of the outer bracket 153 may be fixed to one end of the built-in camera recess 15 via a first movable pivot 152. One end of the outer bracket 153 can be connected to one end of the inner bracket 155 via the second movable pivot 154. The outer bracket 153 may have an opening for the inner bracket 155. When the movable camera bracket is retracted, the inner bracket 155 is configured to fit into the inner bracket 155 opening, further saving space occupied by the movable camera bracket. The inner bracket 155 may have a camera mounting hole. The shape of the miniature camera 157 can match the shape of the camera mounting hole. The miniature camera 157 can be connected to the inner bracket 155 via the third movable pivot 156. The outer bracket 153 can be configured to rotate via the first movable pivot 152. The inner bracket 155 can be configured to rotate via the second movable pivot 154. The miniature camera 157 can be configured to rotate about the third movable pivot 156 as its central axis. All of the above rotation methods can be triggered manually or by other feasible physical means, and are not specifically limited here. The aforementioned miniature camera 157 is configured to communicate with the aforementioned main control unit so that the aforementioned recording device can acquire the function of recording video.
[0112] The aforementioned miniature camera and related structure, as an inventive point of this application, solves the technical problem that "most recording devices lack video recording functionality, and devices with recording functionality have a single lens angle." The specific factors leading to the lack of video recording functionality in most recording devices and the single lens angle of devices with recording functionality are as follows: most voice recorders cannot connect to other electronic devices, therefore, to simplify their structure, most voice recorders typically do not include a camera unit, resulting in a relatively simple recording function and low versatility. Furthermore, recording devices with camera functionality usually have difficult-to-adjust lens angles, leading to poor flexibility. Solving these factors would add video recording functionality to the recording device, improving its versatility and flexibility. To achieve this effect, this application also provides a recording device that works in conjunction with a movable stand and a miniature camera. On one hand, the movable stand and miniature camera are small and foldable, occupying minimal space and not affecting the portability of the recording device. On the other hand, the movable stand can be moved and adjusted in multiple stages, offering high flexibility. Thus, without affecting the portability of the recording device, a flexible camera structure is provided, improving its versatility and flexibility.
[0113] In some embodiments, the recording system includes a recording device as described in Figures 1-7 and a target terminal connected to the recording device, the target terminal being used to display a recording request. The recording device may include a main control unit, a power management unit, a housing 1, a microphone assembly, and a communication unit as described in the embodiments corresponding to Figures 1-7. The target terminal may include, but is not limited to, mobile phones, tablets, and computers. The connection method may include, but is not limited to, wired and wireless connections, wherein the wireless connection method may include, but is not limited to, Bluetooth connection, Wi-Fi connection, 5G / 4G connection, and NFC connection. The wired connection method may be a connection via a USB data cable. The recording request may include, but is not limited to, start, end, pause, and mark.
[0114] Some embodiments of this application provide a recording system that offers high-quality audio capture and wide compatibility. Specifically, the poor audio quality and limited compatibility of most recording systems stem from the fact that most recording devices are built into common smart devices, whose design is easily affected by the device's structure, resulting in poor audio recording quality. Furthermore, most recording devices can only be connected to the target terminal via wired or wireless connections, limiting their compatibility. Therefore, some embodiments of this application provide a recording system comprising a recording device as described in Figures 1-7 and a target terminal connected to the recording device, the target terminal being used to display a recording request. On one hand, the recording system improves the quality of the captured audio by using an external recording device, thus preventing the microphone array from being affected by the target terminal's structure. On the other hand, the recording device can be connected to the target terminal in various ways, broadening its application range. Thus, a recording system with high-quality audio capture and wide compatibility is provided, enhancing the user experience.
[0115] Referring to Figures 8 to 10, this application also provides a multimodal control and management method based on a recording device. The recording device includes the recording apparatus and recording system described above. Figure 8 is a flowchart 100 of some embodiments of the multimodal control and management method based on a recording device according to this application. This multimodal control and management method based on a recording device includes:
[0116] Step 101: In response to the recording device detected by the communication connection starting the recording operation, the first recording content is acquired.
[0117] In some embodiments, the execution entity of the multimodal control and management method based on the recording device can establish a communication connection with the recording device via a wired or wireless connection. The execution entity can be an electronic device (such as a mobile phone, tablet, or computer) connected to the recording device. The recording device can be an external recording device including a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit. The main control unit can be a digital signal processor (DSP) or an application-specific integrated circuit (ASIC), without specific limitations. The main control unit can be used to connect and manage various hardware modules of the recording device, such as sensors, storage devices, and interface circuits. The power management unit can be a rechargeable lithium-ion battery, without specific limitations. The power management unit can be used to supply power to the recording device. The housing can be a housing used to enclose the internal components of the recording device, with an overall thickness of 2.5mm to 4mm. The structure of the housing can be, but is not limited to, plate-shaped or ring-shaped, as long as it can enclose the internal components without interfering with their functionality, without specific limitations. The microphone assembly can include at least one microphone. For example, each microphone can have a diameter of 2mm and a length of approximately 5mm, and can be manufactured using MEMS (Micro-Electro-Mechanical Systems) technology. The microphones can be omnidirectional microphones; for example, each omnidirectional microphone can be 3mm in size, and can be arranged in two rows and two columns within the housing to form an array for 360° sound capture, without specific limitations. The microphone assembly can be used to collect external sounds. The communication unit can be used to communicate with the execution entity. The first recorded content can be audio. The connection method of the communication unit can include, but is not limited to, NFC connection, 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra-wideband) connection, and other currently known or future wireless connection methods. In practice, after establishing the communication connection, the recording device can start recording by pressing the corresponding button (start / stop recording), and simultaneously, NFC can be enabled to trigger a pop-up window of the corresponding app of the execution entity. The recording device can control the execution entity to enter the recording state and acquire the recorded content within the execution entity. Taking NFC connectivity as an example, when NFC devices establish a connection via the NFC communication protocol (such as ISO / IEC 14443), they transmit trigger commands. Preset commands or data are written to the NFC tag. After NFC is triggered, the relevant interface is called through the application layer (such as a video recording application in Android / iOS systems) to start the recording function.Specific implementation methods: The Android system listens for NFC events through the NfcAdapter API. When the mobile terminal receives a trigger signal from an NFC tag on an electronic device, the system performs a specific operation. iOS implementation: The iOS system uses the Core NFC framework to scan and read NFC tags. When an iPhone or iPad touches an NFC tag, the application reads the tag data and triggers a recording reminder.
[0118] In some embodiments of this application, the aforementioned executing entity may also generate recording minutes information based on the first recording content. The recording minutes information may be a text version or a summary of the recorded content. For example, if a meeting was recorded, the recording minutes information may include the meeting's theme, participants, meeting content, meeting time, and meeting date.
[0119] The aforementioned generation of recorded minutes can be achieved using a pre-set text algorithm within the recording device. In practice, the executing entity first uses a speech recognition algorithm to convert the recorded content into text, obtaining a text version of the recorded content. As an example, speech recognition algorithms can include, but are not limited to, Hidden Markov Models (HMMs) based on deep learning, Recurrent Neural Network Transducers (RNNs). Taking a meeting recording as an example, a Hidden Markov Model (HMM) based on deep learning is used to convert the speeches of all parties in the meeting into text. Suppose that at the beginning of the meeting, the host says, "Today's meeting topic is discussing the plan for the next quarter." The speech recognition algorithm accurately recognizes this speech and converts it into text, providing the basic text for the subsequent generation of recorded minutes.
[0120] Then, keywords can be extracted from the recorded text to obtain a keyword set. This keyword set helps identify the core points of the recorded content. As an example, methods for extracting keywords from text can include, but are not limited to, methods based on Term Frequency-Inverse Document Frequency (TF-IDF) and Text Rank algorithms. For instance, using the TF-IDF algorithm, words like "next quarter" and "plan" can be identified as keywords due to their high frequency and relative uniqueness in the meeting text. These keywords will become key elements in generating the minutes information.
[0121] Next, the keyword set of the recorded content can be structured to obtain a set of key information, such as extracting the speaker, time, and key points. Structured processing methods can include Named Entity Recognition (NER) technology, which can identify entities such as names of people, organizations, and times in the text. For example, in the meeting text, NER technology can be used to identify information such as "Speaker: Mr. Li" and "Time: 9:00 AM, October 15, 2024". For key points, by analyzing sentence structure and keywords, key information such as "Mr. Li proposed to expand the southern market and increase advertising efforts in the next quarter" can be extracted and presented in a structured form, such as "Key Point - Speaker: Mr. Li, Content: Expand the southern market and increase advertising efforts in the next quarter".
[0122] Finally, the above key information set can be integrated to obtain the recording minutes information. As an example, the recording minutes information generation algorithm can include, but is not limited to, Text Rank or Sequence-to-Sequence (Seq2Seq) models. For example, using a Sequence-to-Sequence (Seq2Seq) model, sentences containing key information can be selected from the meeting text to generate the following recording minutes information: "At 9:00 AM on October 15, 2024, General Manager Li chaired a meeting to discuss the sales plan for the next quarter, proposing to expand the southern market and increase advertising investment."
[0123] Step 102: Set the first recording content using the recording device.
[0124] In some embodiments, the executing entity can perform setting operations on the first recorded content through the recording device. These setting operations may include, but are not limited to, marking, segmenting, and generating a subject summary. In practice, these setting operations can be implemented through interaction between the recording device and a connected terminal. The interaction method may be controlling the connected terminal to display corresponding operation controls, such as buttons for corresponding functions.
[0125] In some optional embodiments, the execution entity may perform a marking process on the first recorded content through the following steps to obtain the marked first recorded content:
[0126] The first step is to parse the preset target information in the first recorded content to obtain the parsing results. This preset target information may include, but is not limited to, time, location, names, and topic, and can be adjusted as needed without specific limitations. In practice, parsing can also be achieved using the text algorithm described above. The parsing results can be a summary of the key points of the corresponding recorded content, such as the content theme, main participants, and key content elements.
[0127] The second step is to mark the first recorded content based on the above analysis results. In practice, this marking process can be performed in the background.
[0128] In some optional embodiments, the execution entity may perform the following steps to mark the first recorded content based on the parsing results:
[0129] The first step involves marking the preset target in the first recorded content in response to the determination that the parsing result satisfies the marking trigger condition corresponding to the preset target. The marking trigger condition for satisfying the preset target can be: if the parsing result identifies a target person or keyword, then the words spoken by the target person or keyword are marked. For example, if someone starts speaking, the marking process automatically begins for what that person says. The preset target can be a user or an event. For example, if the preset target is a user, the marking trigger condition can be that the user starts speaking. If the preset target is an event, the marking trigger condition can be that the event occurs.
[0130] The second step involves stopping the marking process for the preset target in the first recorded content once the parsing result is determined to satisfy the marking stop condition corresponding to the preset target. The marking stop condition for the preset target can be the identification of another keyword or another operation of the target person. For example, if the preset target is a user, the marking stop condition could be that the user stops speaking. If the preset target is an event, the marking stop condition could be that the event stops.
[0131] In some embodiments of this application, the aforementioned execution entity may also generate recording summary information based on the first recorded content marked with tags.
[0132] In some optional implementations of the embodiments, the execution entity may perform marking processing on the first recorded content based on the parsing results through the following steps:
[0133] The first step is to classify the above analysis results into different levels. In practice, this classification process can be done by comparing the importance of the recorded content with keywords in a preset standard. The specific classification method can be pre-set. For example, when the recording device analyzes the recorded content, if it identifies these preset key contents, such as "the definition and formula of derivatives" in a mathematics course, it will mark the analysis result as high-level. For supplementary examples and extended extracurricular knowledge, such as applications of derivatives in economics, they can be marked as medium-level. Non-critical information, such as routine greetings in class and simple transitional statements, will be marked as low-level.
[0134] The second step is to mark the first recorded content based on the above hierarchical analysis results. In practice, for example, when recording a math lesson, if key content such as the definition of the derivative is detected, it is highlighted with a preset color on the timeline of the recorded content, such as being marked in red on the corresponding timeline. Different levels of recorded content can be marked with different colors. Here, there are no restrictions on the specific form of the marking.
[0135] In some embodiments of this application, the execution entity may also generate second recording content based on the first recording content and the obtained screen recording content. The second recording content may include, but is not limited to, screen recordings, audio recordings, recording information summaries, and fused recording content.
[0136] In optional embodiments of some examples, the execution entity can perform fusion processing on the first recorded content and the screen recorded content to obtain a second recorded content. The first recorded content includes audio from the environment where the recording device is located, and the second recorded content includes the screen recorded content and the corresponding audio from that environment. In practice, the video and audio of the second recorded content need to be precisely matched in time to ensure accurate correspondence between the audio and the corresponding screen image. This fusion processing can be achieved by recording and managing the timestamps of the audio and video during the recording process. During fusion processing, the system arranges and combines the audio and screen images in chronological order based on these timestamps, ensuring that the sound and image are not misaligned or out of sync when playing the recorded content. Furthermore, to achieve fusion, format matching or conversion is required. For example, audio may use common formats such as MP3 and WAV, while screen recorded content may exist in video formats such as MP4 and AVI. During fusion processing, the system detects the data formats of both and converts one or both to a compatible format according to the requirements of the target recorded content. This may involve encoding and decoding operations to ensure the integrity and quality of the data during the conversion process.
[0137] Step 103: In response to the detection of a preset operation by the recording device, switch to different recording modes.
[0138] In some embodiments, the execution entity may switch between different recording modes in response to detecting a preset operation through the recording device. The preset operation may be a user clicking a corresponding control, such as "Enter Marking Mode". The recording modes may include, but are not limited to, marking mode, meeting mode, and reply mode.
[0139] In some alternative implementations of embodiments, the execution entity may switch to unmarked recording mode in response to the recording device detecting an end-of-marking operation. In practice, the end-of-marking operation can be implemented by the user clicking the corresponding control in the pop-up window of the marking mode. For example, clicking the "Stop Marking" button in the pop-up window of the marking mode, or pressing the corresponding button on the recording device, such as the stop button.
[0140] In some embodiments of this application, the aforementioned execution entity may further perform the following steps:
[0141] The first step is to acquire audio using the audio acquisition unit included in the recording device. In practice, the audio acquisition unit may include a microphone array, and when recording begins, the recording device can acquire ambient audio through the microphone array.
[0142] The second step involves preprocessing the audio acquired by the audio acquisition unit to obtain the processed recorded audio. This preprocessing may include, but is not limited to, format conversion and noise reduction. Noise reduction can be achieved using noise reduction algorithms, including but not limited to, spectral subtraction and Wiener filtering.
[0143] In some embodiments of this application, the execution entity may also display a recording pop-up window in response to detecting proximity information of the recording device. Both the execution entity and the recording device may include an NFC unit. The NFC unit in the recording device may be located at a preset position for cooperation with the NFC unit in the execution entity. For example, the preset position may be the center position of the recording device on the side closest to the execution entity. The NFC unit in the execution entity and the NFC unit in the recording device can connect through proximity. The proximity information may indicate that the NFC unit in the recording device and the NFC unit in the execution entity are close and have successfully connected. The recording pop-up window may be a pop-up window triggered after the recording device is connected in proximity, used to configure recording items. The recording pop-up window displays at least one recording setting control. At least one recording setting control may include, but is not limited to: a recording control, a synchronous recording and screen recording control, and a synchronous recording and meeting recording control. The recording control can be used to start recording. The synchronous recording and screen recording control can be used to synchronously record and record the content displayed on the execution entity's screen. The synchronous recording and meeting recording control can be used to synchronously record and record the content of an online meeting. Therefore, a pop-up window for configuring recording items can be triggered by a convenient approach or contact operation, simplifying the connection operation between the terminal and the recording device, and allowing users to directly set recording items in the recording pop-up window, simplifying user operation and improving user experience.
[0144] The above embodiments of this application have the following beneficial effects: The multimodal control and management method based on a recording device provided by this application can offer a method for diversified processing of recorded content based on an external recording device. Specifically, the reason why most current methods for managing recorded content are limited to recording and generating content is that users cannot perform further operations on the recorded content, such as marking and organizing. Therefore, some embodiments of this application's multimodal recording method based on a recording device first acquires first recorded content in response to the recording device detecting a start recording operation via a communication connection. This yields preliminary recorded content. Then, the first recorded content is set using the recording device. This allows the user to perform preset operations on the recorded content. Finally, in response to the recording device detecting a preset operation, different recording modes are switched. This process can be achieved by connecting the external recording device to a smart device, and then processing and manipulating the recorded content through the interaction between the smart device and the external recording device. This provides a method for diversified processing of recorded content, improving the user experience.
[0145] Please continue to refer to Figure 9, which shows flowcharts of some other embodiments of the multimodal control management method based on a recording device. The flowchart 200 of this multimodal control management method based on a recording device includes:
[0146] Step 201: In response to the recording device detected by the communication connection starting the recording operation, the first recording content is acquired.
[0147] Step 202: Set the first recording content using the recording device.
[0148] Step 203: In response to the detection of a preset operation by the recording device, switch to different recording modes.
[0149] In some embodiments, the specific implementation of steps 201-203 and the resulting technical effects
[0150] Steps 101-103 in the embodiments corresponding to Figure 8 can be referred to, and will not be repeated here.
[0151] Step 204: In response to the determination that the preset recording start conditions are met, a recording pop-up window is displayed.
[0152] In some embodiments, the executing entity of the multimodal control and management method based on the recording device described above may display a recording pop-up window in response to determining that a preset recording start condition is met. The recording pop-up window displays a screen recording authorization control. The preset recording start condition may be that the executing entity receives a signal allowing recording, thus enabling the relevant screen recording function to start. The screen recording authorization control may be an operation interface displayed in the recording pop-up window for the user to decide whether to allow screen recording; for example, the executing entity device may display a pop-up window with a "Allow recording?" option.
[0153] Step 205: In response to the detection of a confirmation operation applied to the screen recording authorization control, screen recording is performed to obtain the screen recording content.
[0154] In some embodiments, the aforementioned executing entity may, in response to detecting a confirmation operation applied to the screen recording authorization control, perform screen recording to obtain the screen recording content. This recording content includes not only visual elements such as static images, text, and graphics displayed on the screen, but may also include, depending on the settings, recording system sounds and ambient sounds collected by the microphone, such as video playback sounds and user explanations. In practice, the screen recording can record the visible area within the screen displayed to the currently executing entity, covering all open application windows, desktop icons, pop-up notifications, etc. If the device supports split-screen functionality, the content displayed in the split screen is also within the recording scope.
[0155] Step 206: Configure the screen recording content.
[0156] In some embodiments, the aforementioned executing entity may perform setting operations on the screen recording content. These setting operations may include, but are not limited to, marking and generating real-time subtitles, etc.
[0157] In some optional embodiments, the execution entity can respond to a marking operation detected by the recording device to mark the first recorded content, resulting in marked first recorded content. The marking operation can be a user instructing the recording device to specially mark the currently recorded content. When the execution entity is connected to the recording device, the specific method can be the recording device instructing the execution entity to enter marking mode. The user's interaction with the recording device through the execution entity can include, but is not limited to: touch gesture interaction (e.g., tapping the screen with two fingers; the recording device receives the gesture signal and immediately adds a special mark to the currently recorded content) and voice command interaction (e.g., if the recording device supports voice recognition, the user speaks a specific command, such as "mark the current content," and the device recognizes it and immediately marks the currently recorded content). The special marking of the currently recorded content can include, but is not limited to: pausing, marking, and recording real-time subtitles (which can be generated using the voice recognition algorithm). The marking format can include, but is not limited to: underlines and notes. The marked content can be updated in real-time in the recording minutes information. The marking mode can be entered by operating relevant buttons on the recording device.
[0158] In some embodiments of this application, the executing entity may further store the second recorded content to the recording device and / or a cloud server. The storage device may be an SD card of the recording device. In practice, storing the recorded content in the recording device can reduce the memory pressure on the device where the executing entity is located.
[0159] In some embodiments of this application, the aforementioned execution entity may further perform the following steps:
[0160] The first step is to display a preset meeting recording pop-up window corresponding to the target meeting, in response to the meeting meeting's recording start conditions being met. The recording start conditions can be that the meeting is about to start or has already started. The meeting recording pop-up window displays a screen recording authorization control. The meeting recording pop-up window can be an automatically generated half-screen pop-up window. The meeting recording pop-up window can also be an automatically generated half-screen pop-up window that includes various recording functions (such as pause, start, and end).
[0161] The second step involves detecting a selection operation on the screen recording authorization control displayed on the device, and then recording the target meeting to obtain the meeting recording content. This meeting recording content refers to all information related to the meeting recorded during the screen recording process. This information includes content displayed on the meeting software interface, such as video feeds of participants, shared documents, presentations, chat messages, and whiteboard writing. It may also include audio from the meeting, such as participants' speeches and discussions. It's important to note that simple screen recording differs significantly from the meeting recording described above. For example, simple screen recording has a broad scope, covering all time periods on the screen and various software content, suitable for diverse scenarios such as demonstrations. The meeting recording described above, however, can be recorded separately by recognizing the meeting window of the relevant app. The selection operation can be a user's action on the relevant controls displayed on the smart device, such as clicking the "Start Recording Meeting" button. In practice, the recording device can record meeting-related content by recognizing the content currently displayed on the screen.
[0162] The third step is to identify the above meeting recordings as screen recordings.
[0163] In some embodiments of this application, the aforementioned execution entity may further perform the following steps:
[0164] The first step involves detecting the presence of open settings software content in the screen recording, and displaying a recording pop-up window corresponding to that settings software content. This settings software content can be information displayed, running, or processed by the settings software, including but not limited to text, images, videos, operation menus displayed on the settings software interface, and data generated during background operation. For example, if a browser is set as the settings software, its content could be information about the webpage the user is browsing. If the settings software is office software, the content could be a document or spreadsheet being edited. When the screen recording content is detected, if it contains this type of information from the settings software, subsequent actions are triggered. If it is set as email software, the content could be the email content displayed on the screen. The recording pop-up window displays a screen recording authorization control.
[0165] The second step involves detecting a selection operation on the screen recording authorization control, and then recording the content of the aforementioned software to obtain the recorded content. In practice, users can slide the screen of the smart device to display the content of the aforementioned software and obtain the corresponding screen recording.
[0166] In some embodiments of this application, the aforementioned execution entity may further perform the following steps:
[0167] The first step is to parse the recorded content of the corresponding software to obtain the parsing results. These parsing results can include the main body, key figures, and key summaries of the software content. For example, the recognition model can be Natural Language Processing (NLP) technology. For instance, NLP technology can be used to recognize email content, and deep learning-based NER models, such as BERT-BiLSTM-CRF, can be used to identify key entities in the email, such as product names, customer needs, and dates.
[0168] The second step involves performing preset operations on the software content based on the analysis results described above, using the recording device. These preset operations may include, but are not limited to, replying, summarizing, and marking.
[0169] In some optional embodiments, the execution entity may perform preset operations on the preset software content through the recording device based on the content parsing results using the following steps:
[0170] The first step involves generating an email reply based on the content parsing results of the software used to detect a reply action (specifically, an email type). This reply action could be the detection of a user clicking the cursor in the reply box. In practice, the reply content can be pre-defined email reply content or content generated by a pre-trained language model based on the email content. For example, the pre-trained language model could be GPT-3.
[0171] The second step is to input the email reply content using the aforementioned recording device. In practice, the recording device can also provide a virtual Bluetooth mobile phone keyboard. When the user needs to reply to an email, they only need to move the cursor to the set position to automatically input text.
[0172] In some optional embodiments, the execution entity may perform preset operations on the preset software content through the recording device based on the content parsing results using the following steps:
[0173] The first step involves generating a message reply based on the content parsing results of the software, which is configured as a chat application, in response to the detected reply action. This reply action could be triggered by the user clicking the cursor in the reply box or by the user verbally instructing the recording device to perform an action, such as saying directly to the recording device, "Please reply to this email for me."
[0174] The second step is to input the message reply content using the aforementioned recording device. The input principle and process are the same as those for inputting the email reply content.
[0175] Step 207: Audio is acquired through the audio acquisition unit included in the recording device.
[0176] In some embodiments, the execution entity of the above-described multimodal control and management method based on the recording device can acquire audio through an audio acquisition unit included in the recording device. The audio acquisition unit can be a microphone array within the recording device. The microphone array includes at least one microphone. In practice, the recording device can acquire audio through the audio acquisition unit.
[0177] Step 208: Preprocess the audio obtained by the audio acquisition unit to obtain the processed recorded audio.
[0178] In some embodiments, the executing entity of the above-described multimodal control and management method based on the recording device can preprocess the audio obtained by the audio acquisition unit to obtain the processed recorded audio.
[0179] In some optional embodiments, the execution entity may preprocess the audio obtained by the audio acquisition unit to obtain processed recorded audio through the following steps:
[0180] The first step is to denoise the audio acquired by the audio acquisition unit to obtain denoised recorded audio. Specifically, a preset denoising algorithm can be used to remove background noise from the recorded speech to obtain denoised recorded audio.
[0181] The second step involves adjusting the volume of the denoised audio recording in response to the preset volume difference condition. This results in a processed audio recording. The preset volume difference condition can be that the volume fluctuation of the same speaker or multiple speakers at different times exceeds a preset standard. This standard can be adjusted according to actual needs. For example, the volume difference can be the volume difference between two adjacent sentences or several adjacent words. For instance, if a speaker speaks at a normal volume in one sentence and then suddenly increases the volume in the next due to emotional excitement, a volume difference will occur. When this difference reaches a preset volume difference, such as exceeding 10-15 dB, volume adjustment can be triggered, indicating the need for dynamic range compression. This volume adjustment process can use a preset dynamic balancing algorithm to balance the volume of different sounds in the recording. As an example, the preset dynamic balancing algorithm can include, but is not limited to, simple threshold-based compression algorithms, root mean square (RMS) compression algorithms, or deep learning-based dynamic range compression algorithms. In practice, adjusting the audio volume can make the volume of each part more balanced, avoiding sudden changes in volume to some extent.
[0182] In some optional implementations of certain embodiments, the aforementioned execution entity may perform noise reduction processing on the audio obtained by the aforementioned audio acquisition unit through the following steps to obtain noise-reduced recorded audio:
[0183] The first step involves sampling and quantizing the audio signal to obtain format-adjusted audio. The audio signal is a continuous analog signal, and sampling can be considered as discretizing the continuous audio signal along a time axis. Values are taken from the audio signal at regular time intervals (sampling periods), resulting in a series of discrete sample points. Common sampling frequencies include 44.1kHz and 48kHz; 44.1kHz means 44,100 sample points are collected per second. A higher sampling frequency allows for a more accurate reproduction of the original audio signal after discretization. Quantization maps these continuous values to a finite number of discrete values. The quantization process determines the value range of each sample point and divides it into several quantization levels. In practice, through sampling and quantization, the audio signal is converted from analog to digital form and conforms to specific digital audio format requirements, facilitating subsequent processing and storage.
[0184] The second step involves framing the formatted audio signal to obtain the framed audio. This framing process divides a longer audio signal into several shorter frames. The audio signal within each frame can be considered stationary, facilitating subsequent frequency domain analysis and other processing. In practice, fixed-length frames can be used, with some overlap between frames. For example, each frame can be 20ms long with a 10ms frame shift, meaning there's a 10ms overlap between adjacent frames. This helps avoid signal abrupt changes at frame boundaries.
[0185] The third step is to convert the segmented audio into tensors, resulting in an audio tensor set. In practice, the sample point values of each frame of audio can be used as elements of the tensor. Each frame corresponds to a tensor, and the tensors of all frames are combined to form the audio tensor set, which is then used for subsequent audio processing.
[0186] The fourth step is to group the tensors in the aforementioned audio tensor set to obtain tensor groups. In practice, to better organize and process audio data, the tensors in the audio tensor set can be grouped according to certain rules. Grouping can be based on factors such as audio characteristics and temporal order. For example, based on the frequency characteristics of the audio, tensors with a higher proportion of high-frequency components can be grouped into one group, and tensors with abundant low-frequency components can be grouped into another group. Alternatively, grouping can be based on temporal order, dividing the audio into several segments according to the chronological order, with the audio tensors corresponding to each segment forming a group. This allows for batch processing in the subsequent audio denoising model, improving processing efficiency.
[0187] Fifth, for each tensor group in the above tensor group set, perform the following steps:
[0188] The first sub-step involves inputting the aforementioned tensor set into the encoding layer of a pre-trained audio denoising model to obtain an audio feature set. This audio denoising model includes an encoding layer, a decoding layer, and an output layer. The denoising model can be a neural network model that takes a tensor set as input and denoised recorded audio as output. The encoding layer may include a first convolutional layer (Conv1D) with a kernel size of 3 and a stride of 1, followed by a batch normalization layer, a ReLU activation function layer, a second convolutional layer (Conv1D) with a kernel size of 5 and a stride of 1, and a pooling layer (Max Pooling1D). The first convolutional layer can be used to initially extract audio features. The batch normalization layer can be used to normalize the data output from the convolutional layers. By normalizing each mini-batch of data along the feature dimension, the mean of the data is 0 and the variance is 1, which helps accelerate model training convergence, reduce gradient vanishing or exploding problems, and improve the model's stability and generalization ability. The second convolutional layer described above can be used to further extract features from the data after the first convolution, normalization, and activation processing. By using larger convolutional kernels and more filters, it can capture more complex and global audio features. The pooling layer described above can use max pooling, with a pooling window size of 2 and a stride of 2. The ReLU activation function layer described above can perform a non-linear transformation on the batch-normalized data.
[0189] In practice, the overall function of the aforementioned coding layer is to extract and compress features from the input audio tensor, transforming the original audio data into a more compact and representative feature representation. Through multiple convolutional and pooling operations, different levels of features in the audio data are extracted progressively, from simple local features to complex global features, while reducing data dimensionality and providing effective input for subsequent decoding layers.
[0190] The aforementioned decoding layer may include a first transposed convolutional layer (Transposed Conv1D) with a kernel size of 5 and a stride of 1, a batch normalization layer, a second transposed convolutional layer (Transposed Conv1D) with a kernel size of 3 and a stride of 1, and a third convolutional layer (Conv1D) with a kernel size of 1 and a stride of 1.
[0191] The first transposed convolutional layer described above can be used to correspond to the parameters of the last convolutional layer in the encoding layer to achieve initial recovery of the encoded features. Through inverse convolution, the low-dimensional encoded features can be mapped back to a higher-dimensional space, beginning the reconstruction of the audio's temporal information. The batch normalization layer described above normalizes the data output from the transposed convolutional layer, serving the same purpose as batch normalization in the encoding layer, ensuring the data remains within a suitable range, which helps stabilize the model's training. The second transposed convolutional layer described above can be used to further upsample the data and fuse features, gradually recovering the audio's detailed information. The third convolutional layer described above can be used for final feature integration of the data processed by multiple transposed convolutions, merging multiple feature channels into one channel to obtain a reassembled audio that matches the dimensionality of the original audio data.
[0192] In practice, the aforementioned decoding layer can be used to progressively restore the compact feature representation extracted by the encoding layer to a temporal representation similar to the original audio data. By gradually increasing the data dimensionality through transpose convolution operations, the temporal resolution of the audio is restored. At the same time, by combining convolution and activation functions, features are fused and adjusted to reconstruct the detailed information of the audio, providing the output layer with audio data close to the original audio format.
[0193] The output layer described above can include a fully connected (Dense) layer and a Tanh activation function layer. The fully connected layer contains the same number of neurons as the audio data samples, fully connecting the one-dimensional audio data output from the decoding layer and performing a final weighted summation operation on each sample point. The Tanh activation function layer can apply the Tanh (hyperbolic tangent) activation function, mapping the data output from the fully connected layer to the range of -1 to 1. The Tanh function is suitable for handling the amplitude range of audio signals, normalizing them to conform to a reasonable range of audio values. In practice, the output layer's role is to perform the final adjustment and mapping of the audio data reconstructed by the decoding layer, converting it into a denoised audio segment that meets the expected specifications.
[0194] The second sub-step involves inputting the aforementioned audio feature set into the aforementioned decoding layer to obtain the reassembled audio, wherein the aforementioned decoding layer includes at least one convolutional layer and at least one pooling layer.
[0195] The third sub-step involves outputting the reassembled audio through the output layer to obtain the denoised audio segment.
[0196] The sixth step is to stitch the denoised audio segments together to obtain the denoised recorded audio. In practice, the denoised audio segments can be stitched together in their original time order to obtain the denoised recorded audio.
[0197] The first to sixth steps and related content described above, as an inventive point of this application, solve the technical problem that "most current recording devices have poor noise reduction performance." The factors contributing to the poor noise reduction performance of most current recording devices are often as follows: most recording devices lack a process for further noise processing, making the captured audio easily affected by noise. To achieve this effect, firstly, the captured audio is processed by frame segmentation, cutting the continuous audio signal into short frames with relatively stable characteristics, which helps to more accurately capture noise features in the audio. Simultaneously, through tensor transformation, the framed audio data is converted into a tensor form suitable for deep learning model processing, greatly improving the efficiency and accuracy of data processing. Subsequently, the tensors are grouped and input into a carefully designed noise reduction model. The encoding layer of this noise reduction model can effectively extract key features in the audio and suppress noise, while the decoding layer reconstructs a clear audio signal based on these features, and the output layer further optimizes the audio quality, achieving refined noise reduction for each audio segment. Thus, the quality of audio noise reduction is improved, providing users with a clearer and purer audio recording effect.
[0198] As shown in Figure 9, compared to the description of some embodiments corresponding to Figure 8, the process 200 of the multi-modal control management method based on the recording device in some embodiments of Figure 9 adds a diversified processing flow for the recorded content, providing recording modes to cope with various recording scenarios. The above process also includes pop-up authorization and optimization processing of the recorded audio. This optimization improves the sound acquisition quality and versatility of the above-mentioned multi-modal control management method based on the recording device.
[0199] Referring now to FIG10, a schematic diagram of the structure of an electronic device (computing device 101 as shown in FIG8) 300 suitable for implementing some embodiments of the present application is shown. The electronic device shown in FIG10 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present application.
[0200] As shown in Figure 10, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0201] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 shows electronic device 300 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Each box shown in Figure 10 may represent one device, or multiple devices may be represented as needed.
[0202] According to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this application.
[0203] It should be noted that, in some embodiments of this application, the computer-readable medium described may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0204] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0205] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire first recording content in response to a recording device detected by a communication connection; perform setting operations on the first recording content via the recording device; and switch between different recording modes in response to a preset operation detected by the recording device.
[0206] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0207] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0208] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0209] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.
Claims
1. A recording device, wherein, The recording device includes a main control unit, a power management unit, a housing, a microphone assembly, and a communication unit, wherein... The power management unit and the microphone assembly are disposed within the housing, wherein the microphone assembly is configured to collect audio from different directions outside. The main control unit and the communication unit are disposed within the housing, and the recording device is configured to communicate with the target terminal via the communication unit.
2. The recording apparatus according to claim 1, wherein, The recording device has one of the following structures: card-type or wearable.
3. The recording apparatus according to claim 2, wherein, The recording device has a card-type structure, and the thickness of the card-type recording device is no more than 1 cm.
4. The recording apparatus according to claim 2, wherein, The size of the card-type recording device covers the set size of the target terminal.
5. The recording apparatus according to claim 1, wherein, The recording device is provided with a magnet, which is disposed on a bonding surface on the outer shell for bonding with the target terminal. The magnet is configured to allow the outer shell to be bonded to the target terminal by magnetic attraction.
6. The recording apparatus according to claim 1, wherein, The recording device also includes a data charging cable.
7. The recording apparatus according to claim 6, wherein, The data charging cable includes an external interface and is positioned around the bottom of the housing.
8. The recording apparatus according to claim 6, wherein, The bottom of the housing is provided with a groove, and the data charging connection cable is arranged around the groove.
9. The recording apparatus according to claim 6, wherein, The bottom of the housing is provided with the outlet of the data charging cable, and the data charging cable is designed to be able to be pulled out of the outlet.
10. The recording apparatus according to claim 1, wherein, The outer casing is provided with at least one button.
11. The recording apparatus according to claim 1, wherein, The outer casing is equipped with a touch panel.
12. The recording apparatus according to claim 1, wherein, The touchpad is positioned opposite the recording device and the target terminal, close to each other.
13. The recording apparatus according to claim 1, wherein, The communication unit includes at least one of the following: Bluetooth module, Wi-Fi module, 5G module, 4G module, and NFC module.
14. The recording apparatus according to claim 1, wherein, The recording device also includes an indicator light, which is disposed in a hole on the housing and is communicatively connected to the main control unit.
15. The recording apparatus according to claim 10 or 11, wherein, The at least one button includes a power button and a mode switching button, and the touchpad is used to control the setting operations of the recording device.
16. The recording apparatus according to claim 1, wherein, The recording device is also equipped with a shooting unit, which is configured to acquire external images.
17. The recording apparatus according to claim 16, wherein, The shooting unit is located on top of the recording device.
18. The recording apparatus according to claim 1, wherein, The recording device also includes a vibration unit, which is communicatively connected to the main control unit. The main control unit is configured to control the vibration unit to perform vibration operations.
19. The recording apparatus according to claim 1, wherein, The outer shell has a micro-adhesive pad on the bonding surface for attaching with the target terminal.
20. The recording apparatus according to claim 1, wherein, The microphone assembly includes individual microphones distributed on at least two sides of the housing.
21. The recording apparatus according to claim 1, wherein, The recording device includes a signal processing unit disposed in the housing, and the signal processing unit includes a processing chip.
22. The recording apparatus according to claim 6, wherein, The power management unit is configured to connect to the data charging cable, and the external interface of the data charging cable is configured to connect to a power supply device to supply power and / or charge the power management unit. The power supply device includes at least one of the following: the target terminal and a charging device.
23. The recording apparatus according to claim 1, wherein, The recording device communicates by being attached to the target terminal.
24. A recording system, wherein, Includes a recording device as described in any one of claims 1-23 and a target terminal connected to the recording device, wherein the target terminal is used to display a recording request.
25. A multimodal control and management method based on a recording device, wherein, include: In response to the recording device detected by the communication connection starting the recording operation, the first recording content is acquired; The recording device is used to set the first recording content; In response to the detection of a preset operation by the recording device, different recording modes are switched.
26. The multimodal control and management method according to claim 25, wherein, The method further includes: In response to the determination that the preset recording start conditions are met, a recording pop-up window is displayed, wherein the recording pop-up window displays a screen recording authorization control; In response to detecting a confirmation operation applied to the screen recording authorization control, screen recording is performed to obtain the screen recording content; Configure the screen recording content.
27. The multimodal control and management method according to claim 25, wherein, The method further includes: Audio is acquired through the audio acquisition unit included in the recording device; The audio acquired by the audio acquisition unit is preprocessed to obtain the processed recorded audio.
28. The multimodal control and management method according to claim 25, wherein, The method further includes: Based on the first recorded content, a recording summary is generated.
29. The multimodal control and management method according to claim 25, wherein, The step of setting the first recorded content through the recording device includes: In response to a marking operation detected by the recording device, the first recorded content is marked to obtain the first recorded content with marking.
30. The multimodal control and management method according to claim 29, wherein, The step of switching different recording modes in response to detecting a preset operation through the recording device includes: In response to the recording device detecting the end-of-mark operation, it switches to unmarked recording mode.
31. The multimodal control and management method according to claim 29, wherein, The step of marking the first recorded content to obtain marked first recorded content includes: The preset target information in the first recorded content is parsed to obtain the parsing result; Based on the parsing results, the first recorded content is marked.
32. The multimodal control and management method according to claim 31, wherein, The step of marking the first recorded content based on the parsing result includes: In response to determining that the parsing result satisfies the marking triggering condition corresponding to the preset target, the preset target is marked in the first recorded content; In response to the determination that the obtained parsing result satisfies the marking stop condition corresponding to the preset target, the marking processing of the preset target in the first recorded content is stopped.
33. The multimodal control and management method according to claim 29, wherein, The method further includes: Generate recording summary information based on the first recorded content with tags.
34. The multimodal control and management method according to claim 31, wherein, The step of marking the first recorded content based on the parsing result includes: The parsing results are then subjected to hierarchical processing to obtain hierarchical parsing results; Based on the hierarchical parsing results, the first recorded content is marked.
35. The multimodal control and management method according to claim 26, wherein, The method further includes: Based on the first recorded content and the obtained screen recording content, a second recorded content is generated.
36. The multimodal control and management method according to claim 35, wherein, The method further includes: The second recorded content is stored in the recording device and / or cloud server.
37. The multimodal control and management method according to claim 35, wherein, The step of generating second recording content based on the first recording content, the screen recording content, or the processed recorded audio includes: The first recorded content and the screen recorded content are fused to obtain the second recorded content, wherein the first recorded content includes the audio of the environment in which the recording device is located, and the second recorded content includes the screen recorded content and the audio corresponding to the environment.
38. The multimodal control and management method according to claim 35, wherein, The method further includes: In response to the fulfillment of the recording start conditions of the target meeting, a preset meeting recording pop-up window corresponding to the target meeting is displayed, wherein the meeting recording pop-up window displays a screen recording authorization control; In response to detecting a selection operation on the screen recording authorization control displayed, screen recording is performed on the target meeting to obtain the meeting recording content; The meeting recording content is determined to be screen recording content.
39. The multimodal control and management method according to claim 35, wherein, The method further includes: In response to detecting that there is an opened setting software in the screen recording content, a recording pop-up window corresponding to the setting software content is displayed, wherein the recording pop-up window displays a screen recording authorization control; In response to detecting a selection operation of the screen recording authorization control applied to the display, screen recording is performed on the content of the specified software to obtain the recorded content of the corresponding specified software as the screen recording content.
40. The multimodal control and management method according to claim 39, wherein, The method further includes: The recorded content of the corresponding setting software is parsed to obtain the content parsing result of the corresponding setting software content; Based on the content parsing results, the preset operation is performed on the set software content through the recording device.
41. The multimodal control and management method according to claim 40, wherein, The step of performing preset operations on the set software content through the recording device based on the content parsing result includes: In response to the detection of a reply operation through the setting software, and the setting software being of email type, an email reply content is generated based on the content parsing result; The email reply content is entered through the recording device.
42. The multimodal control and management method according to claim 40, wherein, The step of performing preset operations on the set software content through the recording device based on the content parsing result includes: In response to the detection of a reply operation through the configured software, and the configured software being a chat type, a message reply content is generated based on the content parsing result; The message reply content is input through the recording device.
43. The multimodal control and management method according to claim 27, wherein, The preprocessing of the audio acquired by the audio acquisition unit to obtain the processed recorded audio includes: The audio acquired by the audio acquisition unit is denoised to obtain denoised recorded audio. In response to the denoised recorded audio meeting a preset volume difference condition, the volume of the denoised recorded audio is adjusted to obtain the processed recorded audio.
44. The multimodal control and management method according to claim 25, wherein, The method further includes: In response to the detection of proximity information of the recording device, a recording pop-up window is displayed, wherein at least one recording settings control is displayed in the recording pop-up window.
45. An electronic device, wherein, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal control management method as described in any one of claims 25 to 44.
46. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 25 to 44.
47. A multimodal control and management method based on a recording device, wherein, The recording device includes a recording apparatus or a recording system, and the multimodal control and management method includes: In response to the recording device detected by the communication connection starting the recording operation, the first recording content is acquired; The recording device is used to set the first recording content; In response to the detection of a preset operation by the recording device, different recording modes are switched.