Data processing method, device and system and terminal equipment
By acquiring recording data and marking information and using artificial intelligence models to generate summary information, the problem of unintelligent recording data processing in existing technologies is solved, and efficient and personalized summarization of recording data is achieved.
Patent Information
- Application Number
- CN202510771387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-23
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are unable to intelligently process recording data to meet individual user needs, resulting in users having to spend a lot of time extracting key content from lengthy recordings.
By obtaining recording data and marking information, summary information is generated. The summary information includes abstracts, meeting minutes, to-do lists, etc. The marking information includes the occurrence of human movements or sounds, and is analyzed and summarized using artificial intelligence models.
The generated summary information accurately reflects the user's focus, meets personalized needs, and improves the intelligence and efficiency of recording data processing.
Smart Images

Figure CN120705351A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to a data processing method, a data processing device, a data processing system, and a terminal device. Background Art
[0002] In our daily lives, studies, and work, we sometimes use recording devices to record sounds. However, raw recordings are often lengthy, requiring significant time to decipher and extract the most relevant content. Therefore, intelligently processing recorded data to meet individual user needs has become a pressing issue. Summary of the Invention
[0003] The embodiments of the present application provide a data processing method, apparatus, system, terminal device, and computer-readable storage medium, which can solve the current problem of being unable to intelligently process recorded data to meet individual user needs.
[0004] In a first aspect, an embodiment of the present application provides a data processing method, including: obtaining recording data and marking information, wherein the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device; generating summary information corresponding to the recording data based on the marking information; and outputting the summary information.
[0005] In a possible implementation manner of the first aspect, the summary information includes at least one of the following: a summary, meeting minutes, a to-do list, key issues, and a schedule suggestion.
[0006] In a possible implementation of the first aspect, the marking information includes at least one of the following: human body indication time, human body indication number, the human body indication time includes the human body indication moment and / or the human body indication duration, wherein the human body indication number is used to indicate the number of occurrences of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, the human body indication moment is used to indicate the moment of occurrence of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, and the human body indication duration is used to indicate the duration of human body movements or / and sounds that meet specified conditions during the recording process of the recording device.
[0007] In a possible implementation of the first aspect, the summary information corresponding to the recorded data is generated based on the tag information, including: assigning a first association value to the recorded content portion corresponding to the tag information, the first association value being a first weight or a first attention or a first priority, the first weight being greater than a specified weight, the first attention being greater than a specified attention, and the first priority being higher than a specified priority; and generating the summary information corresponding to the recorded data based on the first association value and the recorded content portion corresponding to the tag information.
[0008] In a possible implementation of the first aspect, the marking information includes a human body indication time. Before generating the summary information corresponding to the recording data based on the first association value and the recording content part corresponding to the marking information, it includes: determining an adjacent time, and the time difference between the adjacent time and the human body indication time is less than or equal to a specified time length value; assigning a second association value to the recording content part corresponding to the adjacent time, the second association value is a second weight or a second attention or a second priority, the second weight is greater than the specified weight, and the second weight is equal to or less than the first weight, the second attention is greater than the specified attention, and the second attention is equal to or less than the first attention, the second priority is higher than the specified priority, and the second priority is equal to or lower than the first priority; correspondingly, generating the summary information corresponding to the recording data based on the first association value and the recording content part corresponding to the marking information includes: generating the summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the marking information, the second association value, and the recording content part corresponding to the adjacent time.
[0009] In a possible implementation of the first aspect, the tag information includes recording device motion time information, and the recording device motion time information is used to indicate the time of the recording device movement controlled by human body motion that meets specified conditions during the recording process of the recording device; correspondingly, generating the summary information corresponding to the recording data based on the tag information includes: determining the interfering action time information based on the sensor data corresponding to the recording device motion time information; filtering the interfering action time information from the recording device motion time information to obtain the filtered recording device motion time information; and generating the summary information corresponding to the recording data based on the filtered recording device motion time information.
[0010] In the second aspect, an embodiment of the present application provides a data processing system, including a recording device and a target terminal, wherein the target terminal is a terminal capable of communicating with the recording device; the recording device is used to: obtain recording data and marking information, the marking information being able to reflect the occurrence of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, and send the recording data and the marking information to the target terminal; the target terminal is used to: receive the recording data and the marking information, generate summary information corresponding to the recording data based on the marking information, and output the summary information.
[0011] In a third aspect, an embodiment of the present application provides a data processing device, including: A data acquisition unit is used to acquire recording data and marking information, wherein the marking information can reflect: the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device; an information generating unit, configured to generate summary information corresponding to the recording data based on the tag information; Information output unit, used to output the summary information In a fourth aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described methods when executing the computer program.
[0012] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.
[0013] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0014] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: in the present application, when the user hears content that he or she considers important, the user can make human body movements or / and language that meet the specified conditions, the terminal device obtains the recording data and marking information, and then generates summary information corresponding to the recording data based on the marking information and outputs the summary information. Since the marking information can reflect: the occurrence of human body movements or / and sounds that meet the specified conditions during the recording process of the recording device, the output summary information can accurately reflect the focus of the user during the recording process, thereby meeting the individual needs of the user, and the output summary information has a stronger correlation with the recording data and has a higher value. As can be seen from the above, the present application can intelligently process the recording data, and the processing steps of the present application are not cumbersome, that is, the present application can efficiently generate relatively accurate summary information. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 This is a flow chart of a data processing method provided in one embodiment of the present application; Figure 2 2 is a schematic diagram of a data processing device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0017] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0018] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0019] It will also be understood that the term "or / and" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0020] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0021] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0022] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized. Example 1
[0023] Figure 1 A flow chart of a data processing method provided in an embodiment of the present application is shown. This data processing method is applied to a terminal device. By way of example and not limitation, the terminal device is a recording device or a target terminal. The target terminal is a terminal capable of communicating with the recording device, for example, a mobile phone, a computer, or a server. The data processing method includes steps S101, S102, and S103, as detailed below: Step S101: Acquire recording data and marking information, where the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device.
[0024] The recording device is a device capable of recording, and the marking information is information recorded by the recording device.
[0025] As an example but not limitation, the recording device is a voice recorder, and the recording data is audio data and / or a transcribed text corresponding to the audio data.
[0026] In some embodiments, the marking information may include at least one of the following: key marking information, recording device movement time information, designated gesture information, and designated voice information, wherein the key marking information can reflect the situation that the designated key of the recording device is pressed by a human body during the recording process, and the recording device movement time information is used to indicate the time during which the recording device is moved by a human body movement that meets the specified conditions during the recording process; the designated gesture information can reflect the occurrence of the designated gesture during the recording process, and the designated gesture is a gesture that can be formed without contacting the recording device, which is equivalent to a designated gesture that can be formed by the human body when separated from the recording device; the designated voice information can reflect the occurrence of the designated voice during the recording process. As can be seen from the above, in this embodiment, the designated key of the recording device is pressed by a human body, the human body movement that can control the movement of the recording device and meets the specified conditions, and the designated gesture are all human body movements that meet the specified conditions, and the designated voice is a human sound that meets the specified conditions, that is, the key marking information, the recording device movement time information, and the designated gesture information can all reflect the occurrence of the human body movement that meets the specified conditions during the recording process, and the designated voice information can reflect the occurrence of the human sound that meets the specified conditions during the recording process. Users can mark importance through body movements and / or sounds.
[0027] Optionally, the marking information includes at least one of the following: human body indication time, human body indication number, the human body indication time includes the human body indication moment and / or the human body indication duration, wherein the human body indication time is used to indicate the time when human body movements or / and sounds that meet specified conditions are generated during the recording process of the recording device, the human body indication number is used to indicate the number of occurrences of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, the human body indication moment is used to indicate the moment when human body movements or / and sounds that meet specified conditions are generated during the recording process of the recording device, the human body indication moment can also be used to indicate the end moment of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, and the human body indication duration is used to indicate the duration of human body movements or / and sounds that meet specified conditions during the recording process of the recording device.
[0028] The human body indication time and the recording time are in the same time system (timing system), that is, in the same time base.
[0029] In this embodiment, the human body indication time can effectively reflect which recording contents are of high importance, and the number of human body indications can effectively reflect the degree of importance. A larger number of indications can play a role of strengthening emphasis, that is, the importance represented by a larger number of indications is greater than that represented by a smaller number of indications.
[0030] In some embodiments, the human body indication time includes the time when the recording device detects human body movements or / and sounds that meet specified conditions during the recording process, which includes a moment (time point) and / or duration (duration). The human body indication count includes the number of times the recording device detects human body movements or / and sounds that meet specified conditions during the recording process. That is, in this embodiment, the time and / or count detected by the recording device can reflect the occurrence of human body movements or / and sounds that meet specified conditions during the recording process.
[0031] In some embodiments, the human body indication times include the number of button presses, the human body indication time includes the button press time, and the button press time includes the button press moment and / or the button press duration.
[0032] By way of example, and not limitation, the recording device is provided with one or more designated buttons. The designated buttons may be physical or virtual buttons. For example, the designated buttons are ergonomically designed physical buttons located on the body of the recording pen, making them convenient for the user to operate with their thumb while holding the pen. During the recording process, the recording device collects and stores audio data. During the recording process, when the user hears content they deem important, they may press the designated button. The recording device is capable of recording the time of key presses, which includes the key press moment (timestamp information) and / or the key press duration. The key press moment includes the press start moment and / or the press end moment. The press start moment is the moment when the designated key begins to be pressed, which can also be interpreted as the moment when the recording device detects that the designated key has been pressed; the press end moment is the moment when the designated key has stopped being pressed, which can also be interpreted as the moment when the recording device detects that the designated key has been released; and the key press duration is the duration of a single press of the designated key.
[0033] In addition, the number of key presses may be the total number of times the designated key is pressed during the entire recording process, or the number of key presses may include the number of times the designated key is pressed corresponding to each consecutive press event. A consecutive press event is when the designated key is pressed at least twice in a row, and consecutive presses indicate that the actual interval between presses of the designated key is less than or equal to the specified interval.
[0034] As an example and not a limitation, assuming that the specified interval time is 0.1 seconds, the time interval between the first time the specified key is pressed and the second time the specified key is pressed is 0.09 seconds, and the time interval between the second time the specified key is pressed and the third time the specified key is pressed is 0.08 seconds, that is, the specified key is pressed three times in a row, and correspondingly, the number of button presses includes 3.
[0035] In this embodiment, the duration of a key press can effectively reflect which recorded content is of high importance, and the number of key presses can effectively reflect the degree of importance. For example, when the number of key presses includes the number of times a designated key is pressed for each consecutive press event, a larger number can serve to enhance emphasis. For example, the importance emphasized by a designated key being pressed three times for a consecutive press event is greater than the importance emphasized by a designated key being pressed two times for a consecutive press event. In other words, a larger number indicates greater importance than a smaller number.
[0036] Optionally, the tag information includes: a type identifier, which may include at least one of the following: a time point type identifier and a time period type identifier. The time point type identifier is an identifier for indicating a time point type, and the time period type identifier is an identifier for indicating a time period type.
[0037] By way of example, and not limitation, each time a user briefly presses a designated key (i.e., each time the recording device detects a short key press event), the recording device records a time point type identifier. In this case, the tag information includes the press start time and the time point type identifier. Thus, each short key press event corresponds to one tag information item (including the press start time and the time point type identifier). The recording device stores each tag information item in a designated log list. Each time a user long presses a designated key (i.e., each time the recording device detects a long key press event), the recording device records a time period type identifier. In this case, the tag information includes the press start time, the press end time, and the time period type identifier. Thus, each long key press event corresponds to one tag information item (including the press start time, the press end time, and the time period type identifier). The recording device stores each tag information item in a designated log list. The recording device associates the recording data with the designated log list. A "short press" indicates that the user presses the designated key for a duration less than or equal to a first duration, while a "long press" indicates that the user presses the designated key for a duration greater than or equal to a second duration, with the first duration being less than the second duration.
[0038] In some embodiments, the terminal device is a target terminal. Accordingly, step S101 includes: the target terminal obtaining the recording data and tag information based on the recording device. That is, before step S101, the recording device transmits the recording data and tag information to the target terminal. For example, the recording device transmits the recording data and tag information to the target terminal via wireless communication technology. Accordingly, step S101 includes: the target terminal obtaining the recording data and tag information transmitted by the recording device. The target terminal then executes subsequent steps S102 and S103.
[0039] By way of example and not limitation, the wireless communication technology is Bluetooth communication technology or Wi-Fi (mobile hotspot) communication technology. The target terminal obtains the recorded data and the designated log list. Because the designated log list records all tag information corresponding to the recorded data, the target terminal obtains the recorded data and all tag information corresponding to the recorded data.
[0040] In some embodiments, the terminal device is a recording device, that is, the recording device executes step S101, step S102 and step S103.
[0041] Step S102: Generate summary information corresponding to the recording data based on the tag information.
[0042] The summary information corresponding to the recorded data is used to represent: a summary result corresponding to the semantics of the recorded data (to be expressed).
[0043] By way of example and not limitation, summary information corresponding to the recorded data is generated based on a designated artificial intelligence model, the tag information, and the recorded data. The input to the designated artificial intelligence model includes the tag information and the recorded data, and the output of the designated artificial intelligence model is the summary information corresponding to the recorded data. The designated artificial intelligence model may be a designated large language model.
[0044] In some embodiments, the step S102 includes: parsing a specified log list, and generating summary information corresponding to the recording data based on the recording data and the parsed tag information.
[0045] In some embodiments, step S101 includes: streaming recording data and tag information, and correspondingly, step S102 includes: streaming analysis and processing of the recording data based on the tag information and the recording data to generate summary information corresponding to the recording data. In other words, this embodiment can implement streaming processing for simultaneous recording and analysis.
[0046] Optionally, the summary information includes at least one of the following: a summary, meeting minutes, a to-do list, key questions, and a schedule suggestion. In other words, the summary information generated in this embodiment more accurately reflects the user's key points and aligns with the main theme of the recording data. This personalized result meets the user's needs and provides greater value.
[0047] Optionally, the recorded data is audio data, and correspondingly, step S102 includes: performing speech recognition processing on the audio data to obtain a transcribed text corresponding to the audio data, marking the text portion corresponding to the marking information in the transcribed text to obtain a marked transcribed text, and generating summary information corresponding to the recorded data based on the marked transcribed text.
[0048] For example, the marking information includes the human body indication time. Accordingly, marking the text portion corresponding to the marking information in the transcribed text includes: marking the text portion corresponding to the human body indication time in the transcribed text. Marking the text portion corresponding to the human body indication time in the transcribed text may include: matching the audio recording time corresponding to the transcribed text with the human body indication time, marking the time position corresponding to the human body indication time on the audio recording timeline, and the transcribed text carrying the marked audio recording timeline, which is equivalent to marking the text portion corresponding to the human body indication time in the transcribed text.
[0049] Optionally, the recording content portion mentioned later in this document may be a text portion; for example, the recording content portion corresponding to the marking information is the text portion corresponding to the marking information.
[0050] In some embodiments, step S102 includes: generating a to-do list based on the tag information.
[0051] Specifically, action items and / or names of responsible persons and / or deadlines are extracted from the portion of the recording content corresponding to the tag information, and a to-do list is generated based on the action items and / or names of responsible persons and / or deadlines.
[0052] It should be noted that the "part of the recorded content corresponding to the marking information" in this specification is used to indicate the content recorded by the recording device during the occurrence / detection of human movements or / and sounds that meet specified conditions, for example, the content recorded by the recording device during the period when a specified button is pressed.
[0053] In some embodiments, step S102 includes: generating meeting minutes based on the marking information.
[0054] Specifically, core topics and / or decision points are extracted from the portion of the recording content corresponding to the marking information, and meeting minutes are generated based on the core topics and / or decision points.
[0055] In some embodiments, step S102 includes: generating a summary based on the tag information.
[0056] Specifically, a summary core is extracted from the portion of the recording content corresponding to the tag information, and a summary is generated based on the summary core.
[0057] In some embodiments, the summary information may be presented in the form of a mind map.
[0058] Optionally, step S102 includes: assigning a first association value to the part of the recorded content corresponding to the tag information, the first association value is a first weight or a first attention or a first priority, the first weight is greater than a specified weight, the first attention is greater than a specified attention, and the first priority is higher than a specified priority; based on the first association value and the part of the recorded content corresponding to the tag information, generating summary information corresponding to the recorded data.
[0059] In this embodiment, the terminal device assigns a greater weight or greater attention or higher priority to the part of the recorded content corresponding to the marking information, so that when the terminal device generates the summary information, it will use the part of the recorded content corresponding to the marking information as a more important reference basis, thereby improving the accuracy of the summary information and better meeting the needs of users.
[0060] In some embodiments, generating summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the tag information includes: generating summary information corresponding to the recording data based only on the first association value and the recording content portion corresponding to the tag information.
[0061] In some embodiments, generating summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the tag information includes: generating summary information corresponding to the recording data based on the first association value and the recording data, that is, generating summary information corresponding to the recording data based on the first association value and the entire recording data, and the entire recording data includes the recording content portion corresponding to the tag information and the unmarked recording content portion.
[0062] Optionally, the marking information includes a human body indication time, and before the summary information corresponding to the recording data is generated based on the first association value and the recording content part corresponding to the marking information, it includes: determining an adjacent time, and the time difference between the adjacent time and the human body indication time is less than or equal to a specified time length value; assigning a second association value to the recording content part corresponding to the adjacent time, the second association value is a second weight or a second attention or a second priority, the second weight is greater than the specified weight, and the second weight is equal to or less than the first weight, the second attention is greater than the specified attention, and the second attention is equal to or less than the first attention, the second priority is higher than the specified priority, and the second priority is equal to or lower than the first priority; correspondingly, the summary information corresponding to the recording data based on the first association value and the recording content part corresponding to the marking information includes: generating the summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the marking information, the second association value and the recording content part corresponding to the adjacent time.
[0063] In actual operation, users may make operational errors. For example, when a user hears important information, he or she does not have time to press the designated button immediately, but is slow to respond for a while. In order to eliminate the negative impact of errors, the terminal device of this embodiment will use the recording content portion corresponding to the marking information and the recording content portion corresponding to the adjacent time as important references when generating summary information, thereby effectively eliminating the negative impact of errors and greatly improving the accuracy of the summary information. Moreover, semantics are often coherent, and the correlation between the recording content portion corresponding to the marking information and the recording content portion corresponding to the adjacent time is also relatively large. Therefore, from this perspective, it can be seen that this embodiment improves the accuracy of the summary information.
[0064] As an example but not limitation, the recording content portion corresponding to the marking information is the text portion corresponding to the marking information; correspondingly, the recording content portion corresponding to the adjacent time is the adjacent text portion of the text portion corresponding to the marking information.
[0065] As an example but not limitation, before generating the summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time, it also includes: assigning a third association value to the recording content part corresponding to the ordinary time, and the third association value is the specified weight or the specified attention or the specified priority; correspondingly, generating the summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time includes: generating the summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value, the recording content part corresponding to the adjacent time, the third association value and the recording content part corresponding to the ordinary time.
[0066] The portion of the recording content corresponding to the normal time is the remaining portion of the recording content in the recording data except the portion of the recording content corresponding to the marking information and the portion of the recording content corresponding to the adjacent time.
[0067] Optionally, the marking information includes recording device motion time information, and the recording device motion time information is used to indicate the time of the recording device movement controlled by human body motion that meets specified conditions during the recording process of the recording device; correspondingly, the step S102 includes: determining the interference action time information based on the sensor data corresponding to the recording device motion time information; filtering the interference action time information from the recording device motion time information to obtain the filtered recording device motion time information; and generating summary information corresponding to the recording data based on the filtered recording device motion time information.
[0068] In which, the recording device is provided with a motion sensor (such as an inertial measurement unit), and correspondingly, the sensor data is: the motion data and / or posture data of the recording device collected by the motion sensor during the recording process of the recording device, and the interference action time information is the time when the interference action occurs.
[0069] In some embodiments, step S101 includes: obtaining recording data, and obtaining tag information, the tag information includes: obtaining total sensor data, the total sensor data is all recording device motion data or / and recording device posture data collected by the motion sensor during the recording process of the recording device, filtering out sensor data that meets predefined characteristics from the total sensor data, and obtaining tag information based on the sensor data that meets the predefined characteristics, the tag information includes recording device motion time information, and the recording device motion time information is used to indicate the collection time corresponding to the sensor data that meets the predefined characteristics. Correspondingly, step S102 includes: determining the interference action time information based on the specified artificial intelligence model and the sensor data corresponding to the recording device motion time information; filtering the interference action time information from the recording device motion time information to obtain the filtered recording device motion time information; and generating summary information corresponding to the recording data based on the specified artificial intelligence model and the filtered recording device motion time information.
[0070] In this embodiment, sensor data that meets predefined characteristics is filtered out from the total sensor data, and interfering action time information is filtered out based on the sensor data corresponding to the motion time information of the recording device based on the specified artificial intelligence model. This is equivalent to performing a secondary screening, which effectively avoids the negative impact of noise caused by relying on sensors for activity detection and significantly improves the accuracy of the summary information.
[0071] Among them, the motion time information of the recording device is used to represent the collection time corresponding to the sensor data that meets the predefined characteristics, that is, the human body movements that meet the specified conditions include human body movements that can cause the recording device to collect the sensor data that meets the predefined characteristics, that is, the collection time corresponding to the sensor data that meets the predefined characteristics can be used to represent the time when the recording device moves controlled by the human body movements that meet the specified conditions during the recording process of the recording device.
[0072] As an example and not a limitation, the interference action time information is the time when the interference action occurs. The interference action may include at least one of the following: rapid rotation, large shaking, repeated tapping, and the interference action meets the predefined characteristics but does not belong to a writing action. This embodiment can filter out the interference action time information that may be caused by the user playing with the voice recorder, tapping with the voice recorder, etc., and the filtered recording device motion time information is used to represent the collection time of the sensor data corresponding to the writing action. That is, in this embodiment, when the user hears content that he or she considers important, the user can perform a writing action on the recording device (such as a voice recorder) to achieve importance marking. The recording device (such as a voice recorder) may have a writing function. The writing action must meet the predefined characteristics.
[0073] In some embodiments, the generating of summary information corresponding to the recorded data based on the specified artificial intelligence model, the tag information and the recorded data includes: judging the effectiveness of each recorded content part corresponding to the tag information based on the specified artificial intelligence model and the contextual scenario of the recorded data, and generating summary information corresponding to the recorded data based on the effectiveness of each recorded content part corresponding to the tag information, each tag information and the recorded data.
[0074] Among them, the designated artificial intelligence model can be used to determine whether the recording content part corresponding to each tag information is a meaningful discussion or key information statement or a task to be performed. If there is a recording content part that is not a meaningful discussion or key information statement or a task to be performed, that is, there is a recording content part that has no practical meaning, then the effectiveness corresponding to the recording content part that has no practical meaning is lower than the effectiveness of the recording content part that has a meaningful discussion or key information statement or a task to be performed.
[0075] By way of example and not limitation, the designated artificial intelligence model may also be capable of optimizing the start and end boundaries marked on the audio data (or transcribed text) based on (contextual) semantics.
[0076] Step S103: output the summary information.
[0077] As an example and not limitation, the summary information is displayed and the user can view, edit, and share the summary information.
[0078] In the present application, when a user hears content that he or she considers important, the user can make human body movements or / and language that meet specified conditions, and the terminal device obtains the recording data and marking information, and then generates summary information corresponding to the recording data based on the marking information and outputs the summary information. Since the marking information can reflect: the occurrence of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, the output summary information can accurately reflect the focus of the user during the recording process, thereby meeting the individual needs of the user, and the output summary information has a stronger correlation with the recording data and has a higher value. As can be seen from the above, the present application can intelligently process the recording data, and the processing steps of the present application are not cumbersome, that is, the present application can efficiently generate relatively accurate summary information.
[0079] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Example 2
[0080] An embodiment of the present application provides a data processing system, which includes a recording device and a target terminal, wherein the target terminal is a terminal capable of communicating with the recording device; The recording device is used to: obtain recording data and marking information, wherein the marking information can reflect: the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device, and send the recording data and the marking information to the target terminal; The target terminal is configured to receive the recording data and the marking information, generate summary information corresponding to the recording data based on the marking information, and output the summary information.
[0081] Optionally, when the target terminal executes the generation of summary information corresponding to the recorded data based on the tag information, it is used to: assign a first association value to the recorded content part corresponding to the tag information, the first association value is a first weight or a first attention or a first priority, the first weight is greater than a specified weight, the first attention is greater than a specified attention, and the first priority is higher than a specified priority; generate summary information corresponding to the recorded data based on the first association value and the recorded content part corresponding to the tag information.
[0082] In some embodiments, the marking information includes a human body indication time. Before the target terminal executes the generation of summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the marking information, the target terminal is further used to: determine an adjacent time, where the time difference between the adjacent time and the human body indication time is less than or equal to a specified time length value; assign a second association value to the recording content portion corresponding to the adjacent time, where the second association value is a second weight or a second attention or a second priority, the second weight is greater than the specified weight, and the second weight is equal to or less than the first weight, the second attention is greater than the specified attention, and the second attention is equal to or less than the first attention, the second priority is higher than the specified priority, and the second priority is equal to or lower than the first priority; correspondingly, when the target terminal executes the generation of summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the marking information, the target terminal is specifically used to: generate summary information corresponding to the recording data based on the first association value, the recording content portion corresponding to the marking information, the second association value, and the recording content portion corresponding to the adjacent time.
[0083] As an example but not limitation, before the target terminal executes the generation of summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time, the target terminal is also used to: assign a third association value to the recording content part corresponding to the normal time, and the third association value is the specified weight or the specified attention or the specified priority; correspondingly, when the target terminal executes the generation of summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time, the target terminal is specifically used to: generate summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value, the recording content part corresponding to the adjacent time, the third association value and the recording content part corresponding to the normal time.
[0084] It should be noted that for technical details not fully described in this embodiment, reference can be made to the data processing methods provided in the various embodiments in the above-mentioned embodiment 1. Example 3
[0085] Corresponding to the data processing method described in the above embodiment, Figure 2 A schematic diagram of a data processing device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0086] The data processing device includes: a data acquisition unit 201 , an information generation unit 202 and an information output unit 203 .
[0087] The data acquisition unit 201 is used to acquire recording data and marking information, where the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device.
[0088] The information generating unit 202 is configured to generate summary information corresponding to the recording data based on the tag information.
[0089] Optionally, when the information generation unit 202 generates summary information corresponding to the recorded data based on the tag information, it is used to: assign a first association value to the recorded content portion corresponding to the tag information, the first association value is a first weight or a first attention or a first priority, the first weight is greater than a specified weight, the first attention is greater than a specified attention, and the first priority is higher than a specified priority; and generate summary information corresponding to the recorded data based on the first association value and the recorded content portion corresponding to the tag information.
[0090] In some embodiments, the marking information includes human body indication time. Before executing the generation of summary information corresponding to the recording data based on the first association value and the recording content part corresponding to the marking information, the information generating unit 202 is further used to: determine the adjacent time, and the time difference between the adjacent time and the human body indication time is less than or equal to the specified time length value; assign a second association value to the recording content part corresponding to the adjacent time, and the second association value is a second weight or a second attention or a second priority, the second weight is greater than the specified weight, and the second weight is equal to or less than the first weight, the second attention is greater than the specified attention, and the second attention is equal to or less than the first attention, the second priority is higher than the specified priority, and the second priority is equal to or lower than the first priority; correspondingly, when executing the generation of summary information corresponding to the recording data based on the first association value and the recording content part corresponding to the marking information, the information generating unit 202 is specifically used to: generate summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the marking information, the second association value, and the recording content part corresponding to the adjacent time.
[0091] As an example but not limitation, before the information generation unit 202 executes the generation of summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time, it is also used to: assign a third association value to the recording content part corresponding to the normal time, and the third association value is the specified weight or the specified attention or the specified priority; correspondingly, when the information generation unit 202 executes the generation of summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value and the recording content part corresponding to the adjacent time, it is specifically used to: generate summary information corresponding to the recording data based on the first association value, the recording content part corresponding to the tag information, the second association value, the recording content part corresponding to the adjacent time, the third association value and the recording content part corresponding to the normal time.
[0092] The information output unit 203 is configured to output the summary information.
[0093] It should be noted that for technical details not fully described in this embodiment, reference can be made to the data processing methods provided in the various embodiments in the above-mentioned embodiment 1. Example 4
[0094] The terminal device in this embodiment includes: at least one processor, memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps of any of the aforementioned data processing method embodiments are implemented. By way of example, and not limitation, the terminal device is a recording device or the aforementioned target terminal.
[0095] When the processor executes the computer program, the steps in the above-mentioned data processing method embodiments are implemented, for example Figure 1 Alternatively, when the processor executes the computer program, the functions of the units in the above-mentioned device embodiments are realized, for example, Figure 2 The functions of units 201 to 203 are shown.
[0096] Those skilled in the art will understand that this embodiment is merely an example of a terminal device and does not constitute a limitation on the terminal device. The terminal device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0097] By way of example and not limitation, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or any conventional processor, etc.
[0098] In some embodiments, the memory may be an internal storage unit of the terminal device, such as a hard drive or memory of the terminal device. In other embodiments, the memory may also be an external storage device of the terminal device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped with the terminal device. Furthermore, the memory may include both the internal storage unit of the terminal device and an external storage device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or is about to be output.
[0099] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0100] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0101] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0102] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0103] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0104] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0105] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0107] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0108] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A data processing method, characterized in that: include: Acquiring recording data and marking information, wherein the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device; generating summary information corresponding to the recording data based on the tag information; The summary information is output.
2. The data processing method according to claim 1, wherein: The summary information includes at least one of the following: a summary, meeting minutes, a to-do list, key issues, and a schedule suggestion.
3. The data processing method according to claim 1, wherein: The marking information includes at least one of the following: human body indication time, human body indication number, the human body indication time includes the human body indication moment and / or the human body indication duration, the human body indication number is used to indicate the number of occurrences of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, the human body indication moment is used to indicate the moment of occurrence of human body movements or / and sounds that meet specified conditions during the recording process of the recording device, and the human body indication duration is used to indicate the duration of human body movements or / and sounds that meet specified conditions during the recording process of the recording device.
4. The data processing method according to claim 1, wherein: Generating summary information corresponding to the recording data based on the tag information includes: Assigning a first association value to the portion of the recorded content corresponding to the tag information, wherein the first association value is a first weight, a first attention level, or a first priority level, wherein the first weight is greater than a specified weight, the first attention level is greater than a specified attention level, and the first priority level is higher than a specified priority level; Summary information corresponding to the recording data is generated based on the first association value and the recording content portion corresponding to the tag information.
5. The data processing method according to claim 4, characterized in that: The marking information includes a human body indication time, and before generating summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the marking information, the method includes: Determine an adjacent time, where a time difference between the adjacent time and the human body indication time is less than or equal to a specified time length value; Assigning a second association value to the portion of the recorded content corresponding to the adjacent time, where the second association value is a second weight, a second attention level, or a second priority level, the second weight is greater than the specified weight, and the second weight is equal to or less than the first weight, the second attention level is greater than the specified attention level, and the second attention level is equal to or less than the first attention level, the second priority level is higher than the specified priority level, and the second priority level is equal to or lower than the first priority level; Correspondingly, generating summary information corresponding to the recording data based on the first association value and the recording content portion corresponding to the tag information includes: Summary information corresponding to the recording data is generated based on the first association value, the portion of the recording content corresponding to the tag information, the second association value, and the portion of the recording content corresponding to the adjacent time.
6. The data processing method according to claim 1, wherein: The tag information includes recording device motion time information, where the recording device motion time information is used to indicate the time during which the recording device is moved by a human body motion that satisfies a specified condition during the recording process of the recording device. Correspondingly, the summary information corresponding to the recording data generated based on the tag information includes: Determining interference action time information based on sensor data corresponding to the motion time information of the recording device; filtering the interfering action time information from the recording device motion time information to obtain filtered recording device motion time information; Summary information corresponding to the recording data is generated based on the filtered recording device movement time information.
7. A data processing system, characterized in that: comprising a recording device and a target terminal, wherein the target terminal is a terminal capable of communicating with the recording device; The recording device is used to: obtain recording data and marking information, wherein the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device, and send the recording data and the marking information to the target terminal; The target terminal is configured to receive the recording data and the marking information, generate summary information corresponding to the recording data based on the marking information, and output the summary information.
8. A data processing device, characterized in that: include: A data acquisition unit, configured to acquire recording data and marking information, wherein the marking information can reflect the occurrence of human body movements and / or sounds that meet specified conditions during the recording process of the recording device; an information generating unit, configured to generate summary information corresponding to the recording data based on the tag information; An information output unit is used to output the summary information.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.