Social software call recording processing method and device based on intelligent terminal, and terminal
By dynamically monitoring the call status of social software on the smart terminal, automatically pop up the recording options and display the status in real time, the problem of lack of recording functions of social software in the existing technology is solved, and convenient recording operations and automated recording file management are realized.
Patent Information
- Application Number
- CN202510460857.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
AI Technical Summary
The social software on existing smart terminals lacks voice call recording functions, which makes it impossible for users to record easily during the call, and the recording operation is cumbersome, making it impossible to view the recording status in real time, affecting the user experience.
By dynamically monitoring the call status of social software on the smart terminal, the recording option will automatically pop up, the recording status will be displayed in real time during the call, and the recording file will be automatically saved after the call is over.
It realizes conveniently starting recording operations during social software calls, displaying recording status in real time, and automatically saving recording files, improving user experience and recording management efficiency.
Smart Images

Figure CN120238602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent terminals, and particularly relates to a method and apparatus for processing call recordings of social software based on an intelligent terminal, an intelligent terminal, and a non-transitory computer-readable storage medium. Background Art
[0002] With the development of technology and the continuous improvement of people's living standards, the use of various intelligent terminals is becoming more and more popular, and various application software with voice call functions will be installed on intelligent terminals. Many mainstream social software on existing intelligent terminals do not have the function of voice call recording. For example, WhatsApp, WeChat, etc. Sometimes it is very inconvenient for users to record calls with social software, and they cannot record the voice information during the social software call in time, which is not convenient for users to use. In addition, the existing recording function often requires users to manually start it, and the recording status cannot be viewed in real time during the call, resulting in a poor user experience. At the same time, the management and subsequent processing of recording files are also relatively cumbersome. Users need to manually perform operations such as converting to text and generating summaries, which increases the complexity of use.
[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and apparatus for processing call recordings of social software based on an intelligent terminal, an intelligent terminal, and a storage medium, which have the advantages of being able to start the recording operation conveniently and quickly during a social software call, displaying the recording status in real time, and automatically saving the recording file after the call ends, thereby improving the user experience; the present invention adds a new function to the intelligent terminal: having the function of recording social software voice calls, which provides convenience for users to use.
[0005] The technical solution adopted by the present invention to solve the problem is as follows:
[0006] The present application provides a method for processing call recordings of social software based on an intelligent terminal, and the technical solution is as follows: when a call request of the social software is detected, control to pop up an optional menu for the recording function; when an operation instruction is received to select to record the current call, perform a recording operation during the call process of the social software, and display the current recording status in real time through a floating window; when it is detected that the social software call ends, control to exit the recording and save the recording file.
[0007] Further, the present application also proposes that when it is necessary to convert the saved recording file to text, receive a convert-to-text operation instruction, control to automatically recognize the selected recording file, convert the selected recording file into text, and automatically generate a summary.
[0008] Furthermore, the present application also proposes to preset a selectable menu for the recording function for recording during a call initiated by a social software.
[0009] Furthermore, the present application also proposes that when a voice signal is detected, it is determined whether the topmost application that initiated the voice signal is a social software for a voice call and whether a voice call is being initiated; when it is detected that the topmost application that initiated the voice signal is a social software for a voice call and a voice call is being initiated, a floating window of the selectable menu for the recording function is controlled to pop up.
[0010] Furthermore, the present application also proposes that when an operation instruction to record the current call is received, the recording operation is automatically started when the social software initiates a call, and the whole process recording or partial recording operation is performed according to the user operation instruction during the call, and the current recording status is displayed in real time through the floating window.
[0011] Furthermore, the present application also proposes that when it is detected that the voice call of the social software ends and a stop playing instruction is received, it is determined whether it is the end of the voice call of the social software that is currently being recorded; if it is the end of the voice call of the social software that is currently being recorded, the recording is controlled to stop, the recording file is saved, and the recording floating window is controlled to exit.
[0012] Furthermore, the present application also proposes to preset three monitors, namely message notification, audio playback type, and task switching, to dynamically monitor the voice call status of the social software in the configured list, so as to implement call recording and the functions of converting the recording into text and generating a summary.
[0013] Furthermore, the present application also proposes a social software call recording processing device based on an intelligent terminal, and the technical solution is as follows: a presetting module, which is used to preset a selectable menu for the recording function for recording during a call initiated by a social software; a recording pop-up control module, which is used to control the pop-up of the selectable menu for the recording function when a call initiation request of the social software is detected; a recording module, which is used to perform a recording operation during the call initiation process of the social software when an operation instruction to record the current call is received, and display the current recording status in real time through the floating window; a recording end and saving module, which is used to control the exit of the recording and save the recording file when it is detected that the call of the social software ends; a text conversion module, which is used to automatically identify the selected recording file and convert the selected recording file into text and automatically generate a summary when a text conversion operation instruction is received for converting the saved recording file into text.
[0014] Furthermore, the present application also provides an intelligent terminal, which includes a memory and one or more programs. One or more of the programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing the above-mentioned method.
[0015] Furthermore, the present application also provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned method.
[0016] As can be seen from the above, the present application provides a method, apparatus, intelligent terminal and non-transitory computer-readable storage medium for processing social software call recordings based on an intelligent terminal. By dynamically monitoring the call status of the social software, the recording option is automatically popped up, flexible recording operations can be realized during the call and the status can be displayed in real time, and the recording file is automatically saved after the call ends. This solves the problems of cumbersome manual operations, inability to view the status in real time and complex subsequent processing in the prior art, and has the advantages of improving the convenience of user operations and the efficiency of recording management. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of a method for processing social software call recordings based on an intelligent terminal provided by an embodiment of the present invention.
[0019] Figure 2 is a schematic flowchart of a recording pop-up window during a voice call of a method for processing social software call recordings based on an intelligent terminal provided by a specific application embodiment of the present invention.
[0020] Figure 3 is a schematic flowchart of the end processing of a voice call of a method for processing social software call recordings based on an intelligent terminal provided by a specific application embodiment of the present invention.
[0021] Figure 4 is a schematic diagram of displaying a call recording floating window of a method for processing social software call recordings based on an intelligent terminal provided by a specific application embodiment of the present invention.
[0022] Figure 5 is a schematic diagram of the recording process of a method for processing social software call recordings based on an intelligent terminal provided by a specific application embodiment of the present invention.
[0023] Figure 6 It is a schematic diagram of converting recording to text and summary for the method of processing social software call recordings based on intelligent terminals provided by specific application embodiments of the present invention.
[0024] Figure 7 It is a block diagram of the principle of the device for processing social software call recordings based on intelligent terminals provided by embodiments of the present invention.
[0025] Figure 8 It is a block diagram of the internal structure principle of the intelligent terminal provided by embodiments of the present invention. Specific Embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer and more explicit, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If this specific posture changes, the directional indications will also change accordingly.
[0028] In the prior art, social software installed in intelligent terminals generally lacks the function of voice call recording. When users use social software for important calls and need to record the call content, they often have to rely on third-party recording devices or manually operate the screen recording function, which makes the recording operation cumbersome and prone to missing key call content. For example, in a business negotiation scenario, users may miss important clause information because they cannot start recording in time, or are forced to interrupt the call and switch to other applications to start the recording function.
[0029] To solve the above problems, the R & D personnel observed the contradiction between users' demand for an instant recording function and the function limitations of social software. Through analysis, it was found that the traditional recording function lacks a linkage mechanism with the call process of social software, resulting in users being unable to quickly trigger the recording operation when the call is initiated. Based on this, the R & D team tried to establish a dynamic association mechanism between the call process and the recording function, and actively provided a recording trigger entry to users at key nodes when the call is established by capturing call state change events of social software.
[0030] Therefore, as Figure 1 shown, this application proposes a method for processing social software call recordings based on intelligent terminals, including the following steps:
[0031] Step S100: When a call request of the social software is detected, control to pop up an optional menu for the recording function;
[0032] Step S200: When an operation instruction to select recording of the current call is received, start recording during the call process in the social software, and display the current recording status in real time through a floating window.
[0033] Step S300: When it is detected that the social software call ends, control to exit the recording and save the recording file.
[0034] Step S500: When it is necessary to convert the saved recording file into text, receive the operation instruction for text conversion, control to automatically recognize the selected recording file, convert the selected recording file into text, and automatically generate a summary.
[0035] Among them, the recording function optional menu refers to the interactive interface triggered according to the call request event detected by the system. Specifically, it can be implemented by calling the system-level window management interface to generate a semi-transparent overlay layer. This interface contains clickable controls for confirming recording and canceling operations. Displaying the current recording status in real time through a floating window means using a visualization component independent of the main application interface. Specifically, it can use the system floating window permission to create a micro panel containing a time progress bar and storage space indication. This component always remains on the top layer of the screen. Saving the recording file means encapsulating and storing the audio stream data in a preset format. Specifically, it can be implemented by dynamically generating a file name containing a timestamp and writing it to the device storage partition.
[0036] Specifically, after detecting the system broadcast of the social software starting a voice call, trigger the permission verification process to obtain the recording permission. When the user selects the recording function, the system automatically associates the audio input source with the file writing process, and at the same time initializes the drawing thread of the floating window. During the call, the audio data is captured in real time through the system audio interface, encoded and then written to the temporary buffer area, and at the same time the timer display of the floating window is updated. When the system monitors the call end event, it automatically triggers the operation of transferring the buffer area data, migrates the temporary file to the user-specified storage path, and releases the occupied recording resources.
[0037] Compared with the prior art, the traditional solution requires the user to manually switch between multiple applications to start the recording function, while this solution realizes the automatic association of the recording function and the call process through the system-level event listening mechanism. In the prior art, there is a problem that the recording is incomplete due to the time difference between the recording operation and the call process. In this solution, it is solved by accurately capturing the call establishment and termination events. In addition, the continuous status feedback of the floating window eliminates the uncertainty of the user about the recording process.
[0038] Through the above technical solution, the present application realizes the instant triggering and precise control of the recording function during the call of the social software. When the call is established, the user can conveniently start recording without interrupting the current operation or switching the application interface. The real-time visual feedback of the recording status enhances the controllability of the operation, and the automatic association mechanism between the recording file and the call process ensures the integrity of the recording content. This solution effectively solves the technical problems of inconvenient recording operation and missing recording content in the traditional social software call scenario.
[0039] The present application further proposes that when it is necessary to convert the saved recording file into text, a text conversion operation instruction is received, the selected recording file is automatically recognized and converted into text, and a summary is automatically generated.
[0040] Among them, the text conversion operation instruction refers to a conversion request triggered by the user clicking a button on the graphical interface or voice input, which can be specifically implemented by a touch event monitoring module or a voice recognition interface, and is used to start the file format parsing and text conversion process.
[0041] Among them, automatic recognition refers to segmenting and semantically analyzing the audio data of the recording file based on a voice feature extraction algorithm, which can be specifically implemented by a voice recognition framework in a deep learning model. For example, spectral features are extracted through a convolutional neural network and matched with a text database.
[0042] Among them, automatically generating a summary refers to extracting keywords and context-related information according to the semantic analysis results to form a structured text, which can be specifically implemented by a text summary generation algorithm in natural language processing, such as a sequence-to-sequence model based on an attention mechanism.
[0043] Specifically, after the user completes the call recording, the target recording file can be selected in the file management interface and the text conversion function can be triggered. The system calls the audio decoding module to perform sampling rate conversion and noise reduction processing on the file, and then converts the audio stream into a text sequence frame by frame through a voice recognition engine. During the conversion process, the system synchronously detects the sentence pause points to divide paragraphs, and generates a summary text including time stamps based on the keyword density and context logic after all conversions are completed. For example, if the call recording contains multiple topic discussions, the summary can be automatically segmented into different topics and the start times can be marked.
[0044] Compared with the prior art, the traditional method requires the user to manually import the recording file into a third-party software for text conversion, and the summary generation depends on manual editing. This solution deeply integrates the recording-to-text function with the call process, reduces the operation steps through automated processing, and at the same time uses semantic analysis technology to realize the instant generation of summaries, avoiding data transmission delays and format compatibility problems caused by switching between multiple tools.
[0045] Through the above technical solution, this application solves the problem that users need to perform additional operations on third-party applications for post-processing of recording files, realizes the integrated operation of call recording and text conversion, enables users to directly obtain structured text and summaries after the call ends, and improves the efficiency of information retrieval and content review.
[0046] This application further proposes that before the step of controlling the exit of recording and saving the recording file when it is detected that the social software call ends, it includes pre-setting a selectable menu for the recording function for recording when the social software starts a call.
[0047] Among them, pre-setting refers to defining parameters and interface layouts for the recording function during the system initialization stage or the user-defined configuration stage. Specifically, it can be implemented by using a user preference setting module or a system default configuration module, and creating a configuration file in the device storage space to record the recording trigger conditions and menu display styles selected by the user. The selectable menu for the recording function refers to a graphical operation interface containing interactive elements such as start recording, pause recording, and stop recording. Specifically, it can be implemented by using a customizable floating control or a system-level pop-up window component, and dynamically generating an interactive layer with touch response function by calling the graphics rendering interface of the operating system.
[0048] Specifically, after the intelligent terminal system starts, the user checks the list of social software through the system settings interface, such as checking application programs such as WeChat and QQ. After the check is completed, the system will generate a corresponding recording permission configuration file and store it in the device memory. During subsequent operations, whenever it is detected that a voice call request is initiated by the configured social software, the system will first read this configuration file and call a preset interface template to generate a semi-transparent floating window containing a circular recording button, a progress bar, and a timer. This floating window can be dragged to any position on the screen, and the transparency can be adjusted to avoid blocking the call interface.
[0049] Compared with the prior art, traditional mobile terminals require users to manually activate the recording function every time they make a call, lacking a pre-configuration mechanism for different social applications. This solution enables the recording operation of common social software to automatically trigger a preset interface by establishing an application whitelist and an interface template preloading mechanism, avoiding the cumbersome operation caused by repeated settings.
[0050] Through the above technical solution, this application effectively solves the technical defect that users need to frequently manually turn on the recording function, and realizes the one-key call of the recording function through the pre-configuration mechanism. When the user makes a call using social software, the system can automatically present a standardized recording operation interface, significantly improving the timeliness and convenience of voice information collection.
[0051] The present application further proposes that the steps for controlling the pop-up of an optional menu for the recording function when a call request is detected in a social software include: when a voice signal is detected, it is determined whether the topmost application that starts the voice signal is a social software for voice calls and whether a voice call is being started; when it is detected that the topmost application that starts the current voice signal is a social software for voice calls and a voice call is being started, a floating window of the optional menu for the recording function is controlled to pop up.
[0052] Among them, voice signal detection refers to collecting an audio signal through the terminal microphone and analyzing whether it is a voice input in a call scenario. Specifically, it can be implemented by combining an audio waveform recognition algorithm with an environmental noise filtering technology to distinguish call voices from background noise. The topmost application judgment refers to obtaining the application information in the foreground running state and identifying whether it is a preset social software. Specifically, it can be implemented by calling the system API interface to read the application process information to ensure that the recording function is triggered only for the target social software. The floating window pop-up refers to displaying a recording operation interface at a specified position on the screen in the form of an overlay layer. Specifically, the system-level window management module can be used to control the rendering and interaction logic of the floating window to provide a non-intrusive user operation entry.
[0053] Specifically, when the user initiates a voice call through a social software, after the terminal microphone collects a voice signal, the signal processing module analyzes whether it is a valid call voice. If it is confirmed that there is a valid voice signal, the application status detection module is further called to determine whether the current topmost application is a preset social software, such as WeChat or WhatsApp. When both conditions are met, the system automatically triggers the floating window generation module to display an operation interface including a recording switch button in the edge area of the screen. This process avoids false triggering through a dual judgment mechanism. For example, no irrelevant recording menu will pop up when the user uses other voice applications.
[0054] Compared with the prior art, the traditional solution usually triggers the recording function only based on the application startup event, which is prone to misoperation in non-call scenarios. However, this solution can accurately determine whether the social software is in the actual call stage by combining voice signal feature recognition and application status detection, effectively reducing the probability of false triggering of the function. In addition, in the prior art, the recording function entry is usually fixedly embedded in the application interface, while this solution uses a floating window form, which can maintain the visibility and convenience of the operation entry in different application scenarios.
[0055] Through the above technical solution, the present application can accurately identify the real call scenario of the social software and dynamically activate the recording function on the premise of ensuring user privacy. This solution solves the problem of misoperation caused by the single triggering mechanism of the recording function in the prior art. At the same time, through the interactive design of the floating window, it improves the timeliness and operation efficiency of function call, and avoids the recording requirement of the user missing the key call period due to interface switching.
[0056] The present application further proposes that when an operation instruction is received to select to record the current call, the recording operation is automatically started when the social software starts the call, and the whole process recording or partial recording operation is performed according to the user operation instruction during the call, and the current recording status is displayed in real time through the floating window.
[0057] Among them, the operation instruction refers to a control signal triggered by the user through touch click, gesture operation or voice input, etc. Specifically, it can be implemented by button trigger or preset gesture recognition module, and is used for the interactive response mechanism to start the recording function. The whole process recording refers to the process of continuously collecting audio data from the establishment to the hanging up of the call. Specifically, it can be realized by calling the system audio interface to ensure the integrity of the recording. The partial recording refers to the segmented collection according to the start / end instruction actively triggered by the user. Specifically, it can be realized by timestamp marking to flexibly intercept the recording content. The floating window refers to a visual control superimposed on the top layer of the application interface. Specifically, it can be constructed by the system window management service and is used to display status parameters such as recording duration and storage space occupancy rate in real time.
[0058] Specifically, after detecting that the user selects the recording operation, the system immediately activates the audio collection thread and establishes a data channel with the call process of the social software. During the call, the floating window continuously displays the waveform diagram and the timer, and at the same time, a segmented control button is built in. When the user clicks the "pause" button, the system generates a time mark and pauses writing the audio stream; clicking the "continue" button again resumes recording from the breakpoint. This dynamic control mechanism allows multiple recording segments to be formed in a single call, and all operation records are finally integrated into a single audio file with segmented marks.
[0059] Compared with the prior art, the traditional recording scheme only supports a single recording mode for the whole call process and cannot dynamically adjust the recording range during the call. Through the interactive design of the floating window, this solution enables the user to adjust the recording strategy in real time according to the importance of the call content. For example, start recording when key information is involved and pause storage during the chatting period, effectively optimizing the utilization rate of the storage space.
[0060] Through the above technical solution, the present application realizes the deep adaptation of the recording operation and the call process, and solves the technical obstacle that existing social software cannot perform segmented recording. Users can flexibly control the recording range according to actual needs, and at the same time timely master the recording status through the visual interface, avoiding the problem of missing important information due to insufficient storage space or operation errors.
[0061] The present application further proposes that when it is detected that the voice call of the social software ends and a stop playback instruction is received, it is judged whether it is the end of the voice call of the social software that is currently being recorded; if it is the end of the voice call of the social software that is currently being recorded, then control the stop of recording, save the recording file, and control the recording floating window to exit.
[0062] Among them, detecting the end of the voice call of the social software means monitoring the change of the call state of the social software, such as listening to the change of the audio stream state or the network connection state of the application process. Specifically, it can be implemented by using the audio focus monitoring interface provided by the operating system to capture the call termination signal. The stop playback instruction is a control signal for terminating the audio output actively triggered by the system or the user. Specifically, it can be implemented by intercepting the pause or stop event of the media player to synchronously terminate the recording operation. Judging whether it is the end of the voice call of the social software that is currently being recorded means associating and matching the call end event with the current recording process. Specifically, it can be implemented by using a comparison mechanism of process identifiers or session numbers to avoid misjudging the audio activities of other applications. Controlling the recording floating window to exit means closing the visual interface elements related to the current recording operation. Specifically, it can be implemented by calling the destruction interface of the window manager to release system resources and end user interaction.
[0063] Specifically, after detecting that the call state of the social software changes from active to terminated, the system will receive a playback termination notification from the media service. At this time, by extracting the identifier of the currently running recording process and matching and verifying it with the application process that triggered the call end. If the identifier matches successfully, then call the audio acquisition interface to immediately terminate the recording thread, and at the same time persistently store the recorded audio data according to the preset path. During this process, the exit operation of the floating window is synchronized with the recording termination event to ensure the immediate feedback of the user interface.
[0064] Compared with the prior art, the traditional solution lacks the process relevance verification when detecting the call end, and may be misjudged due to multiple applications running simultaneously. For example, when there are other voice playback activities in the background, the system may wrongly regard the end event of a non-target call as the recording termination signal. This solution avoids the problem of the recording file being wrongly truncated or incompletely saved by introducing a process identifier matching mechanism to accurately lock the recording session of the target application.
[0065] Through the above technical solution, the present application can accurately identify the termination event of the target social software call, and ensure the integrity and reliability of the recording operation through the process association mechanism. The recording file is immediately saved to the designated storage location after the call ends, and the automatic exit of the floating window reduces the steps of manual operation by the user, effectively preventing the loss of recording data due to operation delays.
[0066] The present application further proposes pre-setting three monitoring methods, namely, message notification, audio playback type and task switching, to dynamically monitor the voice call status of the social software in the configuration list, so as to realize the call recording and the conversion of the recording into text and summary functions.
[0067] Among them, message notification monitoring refers to capturing and analyzing the call request notification of the social software, which can be implemented by using a system broadcast receiver or a notification monitoring service to identify whether the social software initiates a voice call request.
[0068] Among them, audio playback type monitoring refers to the detection of the device audio routing status, which can be achieved through audio focus changes or audio session type judgment, and is used to distinguish between call voices of social software and other media playback behaviors.
[0069] Among them, task switching monitoring refers to monitoring the status of the application switching to the foreground or background, which can be implemented through the activity lifecycle callback or task stack monitoring interface to determine whether the target social software is in a user interaction state.
[0070] The configuration list refers to a list of social software identifiers that are allowed to be monitored, which can be specifically defined by using application package names or process feature matching rules to limit the scope of application of dynamic monitoring.
[0071] Specifically, the message notification monitoring module captures the call request notification of the social software in real time. For example, when WeChat triggers a voice call, the system notification bar will generate a message containing a specific logo; the audio playback type monitoring module synchronously detects whether the device audio channel is switched to call mode, for example, through the audio manager to determine whether the current audio session type is voice communication; the task switching monitoring module continues to track whether the social software is switched to the foreground, such as when the user switches from the background to the WeChat call interface to trigger a status update. The above three monitoring modules work together to dynamically determine whether the social software in the configuration list is in the voice call startup stage. When the preset conditions are met, such as the notification trigger, audio type matching and the application being in the foreground are detected at the same time, the recording function control logic is activated to provide trigger conditions for subsequent recording operations and text conversion functions.
[0072] Compared with the prior art, existing solutions usually rely on a single monitoring method, such as only monitoring notification or audio status, which is prone to misjudgment or missed detection due to complex user operation scenarios. For example, when the user does not enable the notification permission or the call interface is blocked by other applications, the prior art may not be able to trigger the recording function in time. In contrast, this solution combines multi-dimensional status monitoring with comprehensive judgment of messages, audio, and task switching, effectively reducing the probability of false triggering and improving the accuracy of voice call status recognition, thereby ensuring the reliable activation of the recording function in various user scenarios.
[0073] Through the above technical solution, this application can accurately identify the call start status of social software, avoid function failures caused by differences in user operation habits or system permission restrictions, ensure the timely activation of the recording function, and provide complete data support for subsequent voice-to-text conversion and summary generation, solving the problem of untimely or missed triggering of the recording function due to a single monitoring mechanism in the prior art.
[0074] The following further elaborates on the present invention through a specific application example:
[0075] This specific application example adopts three monitoring processes of message notification + audio playback type + task switching to dynamically monitor the voice call status of the APPs in the configured list, thereby realizing call recording and the functions of converting the recording to text and summary.
[0076] A method for processing social software call recording based on an intelligent terminal provided by a specific application example of the present invention specifically includes a recording pop-up window process step and a call end processing process during a voice call; among them, as Figure 2 shown, the recording pop-up window process step during the voice call includes:
[0077] S11: Start;
[0078] S12: Receive a task switching instruction and enter S13;
[0079] S13: Receive a message notification;
[0080] S14: When a message notification is received, enter step S15;
[0081] S15: Determine whether it is a message notification of a matching application. If it is a match, enter S18; if it is not a match, enter step S20;
[0082] S16: When the playback of audio is received, enter S17;
[0083] S17: Determine whether the audio playback type matches. If it matches, enter S18; if it does not match, enter S20;
[0084] S18: Determine whether the currently top - most APP is an APP that matches call recording and is a valid audio playback. If it is valid, enter S19; if it is invalid, enter S20;
[0085] S19: Determine that the currently voice - call - supported APP is in a voice call, pop up a recording floating window, and then enter S20;
[0086] Regarding when making a voice call using a social software, an example of popping up a recording floating window is as follows Figure 4 As shown, display the call recording floating window; the user can select the current voice call recording method, including allowing only when using this application, only this time, or not allowing, three methods.
[0087] As Figure 5 shown, when the user clicks on the floating window, start recording.
[0088] S20: End.
[0089] Among them, as Figure 3 shown, the voice call end - processing flow includes the steps:
[0090] S31: Start, enter S32;
[0091] S32: Monitor the end of audio playback and enter S33;
[0092] S33: Determine whether the audio type matches. If it is, enter S34; if not, enter S37;
[0093] S34: Receive the stop - playback status and enter S35;
[0094] S35: Determine whether it is the APP in the current recording that stops. If it is, enter S36; if not, enter S37;
[0095] S36: Control to stop recording, save the recording file, exit the floating window display, and enter S37;
[0096] In the embodiment of the present invention, when the voice call of the social software ends and exits the recording, save the recording file.
[0097] S37: End.
[0098] In the embodiment of the present invention, as Figure 6 shown, when clicking on the recording notification, enter the call recording interface, and the recording can be converted into text and summary.
[0099] Exemplary device
[0100] As Figure 7As shown in the figure, an embodiment of the present invention provides a social software call recording processing device based on a smart terminal. The device includes:
[0101] A presetting module 310, configured to preset an optional menu for the recording function for recording when a call is initiated by a social software;
[0102] A recording pop-up control module 320, configured to control the pop-up of the optional menu for the recording function when a call request initiated by a social software is detected;
[0103] A recording module 330, configured to perform a recording operation during the call initiation process of the social software when an operation instruction to select recording the current call is received, and display the current recording status in real time through a floating window;
[0104] A recording end and saving module 340, configured to control the exit of the recording and save the recording file when it is detected that the call of the social software ends;
[0105] A text conversion module 350, configured to, when it is necessary to convert the saved recording file into text, receive a text conversion operation instruction, control the automatic recognition of the selected recording file, convert the selected recording file into text, and automatically generate a summary, as described above in detail.
[0106] Among them, the presetting module refers to a program unit for configuring the recording function trigger condition before a call. Specifically, it can be implemented by using an application permission management interface and a system settings database. By writing the list of social software selected by the user and the recording rules, it provides a basic configuration for subsequent recording operations. The recording pop-up control module refers to a functional component for dynamically detecting call requests and triggering an interactive interface. Specifically, it can be implemented by using a system event listener and a graphical interface rendering interface. By capturing the changes in the application process status in real time, it activates the menu display when it detects that a target social software initiates a call. The recording module refers to a functional unit for performing audio acquisition and status feedback. Specifically, it can be implemented by using a system audio input interface and a floating window control. By calling the microphone permission of the terminal and updating the interface elements in real time, it realizes the visibility control of the recording process. The recording end and saving module refers to a processing unit for terminating the recording and storing data. Specifically, it can be implemented by using a file system operation interface and an audio encoder. By listening to the call end event to trigger the file encapsulation process, it converts the original audio data into a standard format for storage. The text conversion module refers to a functional component for converting audio content into text information. Specifically, it can be implemented by using a speech recognition engine and natural language processing algorithms. By parsing the acoustic features of the recording file to generate text content, and automatically generating a summary based on keyword extraction technology.
[0107] Specifically, when a user starts a social software to make a call, the system detects the event through a pre-set monitoring mechanism, and then activates the pop-up operation of the recording function optional menu. After the user chooses to record, the device automatically calls the audio input device of the terminal to collect call data, and at the same time feedbacks real-time information such as recording duration and storage space status through a floating window. During the call, if the user needs to pause or segment the recording, he can send an operation instruction through the floating window control. After the call ends, the device automatically terminates the recording process and encrypts the audio data and stores it in the specified path. When the user needs to check the content of the call, he can choose to input the recording file into the text conversion module, which generates a text record through speech-to-text technology, and automatically extracts key information based on semantic analysis to form a summary.
[0108] Compared with the existing technology, existing social software generally lacks built-in recording functions. Users need to rely on third-party recording software for manual operation, which has problems such as startup delay and interface occlusion. This device realizes system-level integration of recording functions through modular design, automatically activates the interactive menu when a call request is triggered, and avoids the operational burden of manually switching applications. At the same time, the real-time status display of the floating window solves the technical defect that traditional recording software cannot intuitively feedback the recording process, and the integration of the text conversion module further improves the management efficiency of the recording content.
[0109] Through the above technical solution, this application solves the user experience problem caused by the lack of recording function during social software calls, and realizes the automatic triggering and visual control of recording operations. The pre-configured monitoring mechanism ensures the seamless connection between the recording function and the social software, and the real-time status feedback improves the user's control over the recording process, while the speech-to-text function effectively reduces the time cost of later content sorting.
[0110] Based on the above embodiments, the present invention further provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 8 As shown. The smart terminal can be a smart device, including a processor, a memory, a network interface, a display screen, and a database connected through a system bus. Among them, the processor of the smart terminal is used to provide computing and control capabilities. The memory of the smart terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the smart terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for processing call recording of a social software based on a smart terminal is implemented. The database of the smart terminal is used to store a call recording processing program for a social software based on a smart terminal.
[0111] Those skilled in the art will understand that Figure 8The block diagram of the principle shown only shows the block diagram of the partial structure related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0112] In one embodiment, an intelligent terminal is provided, which includes a memory and one or more programs. One or more programs are stored in the memory and are configured to be executed by one or more processors to implement the control function for processing call recordings of social software.
[0113] Among them, the memory refers to the hardware component for storing executable program codes, and specifically can be implemented by a flash memory chip or a solid-state drive. Its function is to save the program file containing the recording control logic.
[0114] Among them, the processor refers to the computing unit that executes program instructions, and specifically can be implemented by a multi-core mobile processor. Its function is to dynamically monitor the call status of social software according to the program code and trigger the recording function.
[0115] Among them, the program refers to a set of codes containing computer instructions, and specifically can be implemented by the Android system service combined with the application layer API call. Its function is to control the execution of the recording process by listening to the call events of social software.
[0116] Specifically, when the program is loaded by the processor, it continuously monitors the call start event of social software on the intelligent terminal. For example, when a WeChat voice call request is detected, the program automatically triggers a system-level pop-up window to display the recording option. After the user selects recording, the program calls the audio acquisition interface to record the call content, and at the same time generates a semi-transparent floating window at the edge of the screen to display the recording duration and storage space status. After the call ends, the program automatically stops recording and encrypts and stores the file in a specified directory. Further, the program can integrate a voice recognition engine to allow the user to convert the saved recording file into text and generate a key information summary.
[0117] In some specific embodiments, the program can dynamically monitor the package name of social software by registering a system broadcast receiver. For example, when a voice call activity of WhatsApp is detected, the program immediately activates the recording function module. In addition, the storage path of the recording file can be set to an encrypted partition independent of the social software data area to prevent third-party applications from illegally reading it.
[0118] Compared with the prior art, the existing solution requires the user to manually start the recording application and switch it to run in the background, which has the problems of cumbersome operation and incomplete recording. This solution realizes the automatic association of call events and the recording function through system-level program integration, and the floating window status feedback avoids the recording interruption caused by the user's misoperation. For example, during a WeChat call, the user can complete the recording control without leaving the current interface, and the recording file is automatically associated with the corresponding contact information.
[0119] Through the above technical solution, this application solves the problem of inconvenient operation caused by the lack of recording function during social software calls, and realizes the automatic matching of the recording process and call events. The user can complete the recording operation without interrupting the call, and the recording file is automatically saved and supports subsequent text conversion processing, improving the efficiency and security of voice information management.
[0120] This application further proposes a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the social software call recording processing method based on the intelligent terminal. The method includes: when detecting a call request initiated by the social software, controlling the pop-up of an optional menu for the recording function; when receiving an operation instruction to select recording the current call, performing a recording operation during the call startup process of the social software, and displaying the current recording status in real time through a floating window; when detecting the end of the social software call, controlling the exit of the recording and saving the recording file; when it is necessary to convert the saved recording file into text, receiving a text conversion operation instruction, controlling the automatic recognition of the selected recording file, and converting the selected recording file into text and automatically generating a summary.
[0121] Among them, the non-transitory computer-readable storage medium refers to a physical carrier for persistently storing computer instructions, and can specifically be implemented by hardware devices such as solid-state drives, flash memory chips, or optical discs. The instructions stored therein can be read and executed by the processor to achieve specific functions. The processor executing the instructions means that the program code in the storage medium is parsed and run through the arithmetic unit of the electronic device, and can specifically be implemented by chip components such as a central processing unit or a graphics processing unit, so as to achieve functions such as call recording, status display, file saving, and text conversion.
[0122] Specifically, when the processor of the electronic device executes the instructions in the storage medium, it first detects the call start request of the social software in real time through a dynamic monitoring mechanism, for example, by identifying message notifications, audio playback types, or task switching states. Once a call request is detected, an instruction to pop up the recording function menu is immediately triggered, and the user can choose whether to start recording. During the call, the recording operation is executed in real time, and the floating window dynamically displays the recording status, for example, updated in the form of a progress bar or time marker. After the call ends, the processor automatically stops recording according to the instruction and saves the file to the specified storage path. If the user needs to convert the recording into text, the processor further calls the speech recognition algorithm to parse the selected file, extract key information, and generate a text summary, for example, through a natural language processing model to achieve semantic analysis and content condensation.
[0123] Compared with the prior art, existing social software generally lacks a built-in call recording function. Users need to rely on third-party applications or manual operations for recording, which has problems such as cumbersome operations, inability to automatically trigger, and real-time feedback. This solution works in coordination with the processor through the instructions in the storage medium to achieve the automated integration of the recording function, without the need to install an additional independent application. At the same time, combined with the floating window status display and text conversion function, a complete recording processing closed-loop is formed.
[0124] Through the above technical solution, this application can effectively solve the problem that users cannot conveniently record during social software calls. By automatically triggering the recording menu, providing real-time status feedback, and performing intelligent processing of recording files, it significantly improves the efficiency and accuracy of voice information recording, while reducing the complexity of manual operations and enhancing the user experience.
[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0126] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing call recording of social software based on an intelligent terminal, characterized in that: include: When a call request is detected from a social software, a pop-up menu for recording function options is displayed; When receiving an operation instruction to record the current call, the social software starts recording the call process and displays the current recording status in real time through a floating window; When it is detected that the social software call ends, the control exits the recording and saves the recording file.
2. The method for processing call recording of social software based on smart terminal according to claim 1, characterized in that: When the social software call is detected to be ended, the step of controlling to exit the recording and saving the recording file includes: When the saved recording file needs to be converted into text, the text conversion operation instruction is received, the selected recording file is automatically identified, the selected recording file is converted into text, and a summary is automatically generated.
3. The method for processing call recording of social software based on smart terminal according to claim 1 is characterized in that: When the social software call is detected to be ended, the step of controlling to exit the recording and saving the recording file includes: An optional menu for the recording function is pre-set for recording calls initiated by social software.
4. The method for processing call recording of social software based on smart terminal according to claim 1, characterized in that: The step of controlling a pop-up recording function optional menu when a call request initiated by the social software is detected comprises: When a voice signal is detected, it is determined whether the top application currently initiating the voice signal is a social software for voice calls, and whether the voice call is being initiated; When it is detected that the pinned application currently initiating the voice signal is a social software for voice calls and the voice call is being initiated, a pop-up recording function optional menu floating window is controlled to be displayed.
5. The method for processing call recording of social software based on smart terminal according to claim 1, characterized in that: The step of receiving an operation instruction to select recording the current call, initiating a recording operation during the call in the social software, and displaying the current recording status in real time through a floating window includes: When an operation instruction is received to record the current call, the recording operation will automatically start when the social software starts the call. During the call, the entire call or part of the call will be recorded according to the user's operation instruction, and the current recording status will be displayed in real time through a floating window.
6. The method for processing call recording of social software based on smart terminal according to claim 1, characterized in that: When the end of the social software call is detected, the step of controlling the recording to exit and saving the recording file includes: When the end of the social software voice call is detected and a stop playback instruction is received, it is determined whether the social software voice call currently being recorded has ended; If the voice call on the social software that is currently being recorded ends, the recording will be stopped, the recording file will be saved, and the recording floating window will be exited.
7. The method for processing call recording of social software based on smart terminal according to claim 1, characterized in that: Before the step of controlling the pop-up recording function optional menu when detecting that the social software initiates a call request, the step also includes: Pre-set three monitoring functions, namely message notification, audio playback type and task switching, to dynamically monitor the voice call status of the social software in the configuration list, so as to realize call recording and convert the recording into text and summary.
8. A social software call recording processing device based on an intelligent terminal, characterized in that: The device comprises: A pre-setting module, used to pre-set an optional menu of a recording function for recording when a call is initiated by a social software; The recording pop-up control module is used to control the pop-up recording function optional menu when a call request initiated by the social software is detected; The recording module is used to start the recording operation of the call process in the social software when receiving the operation instruction to record the current call, and display the current recording status in real time through a floating window; The recording end and save module is used to control the exit of recording and save the recording file when detecting that the social software call is ended; The text conversion module is used to receive text conversion operation instructions, control the automatic recognition of the selected recording file, convert the selected recording file into text, and automatically generate a summary when the saved recording file needs to be converted into text.
9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as described in any one of claims 1 to 7.