Data processing method and apparatus, and storage medium

The server receives the client's parsing instructions, obtains the media file content and sends information to the target model, and processes the output results to generate display information, which solves the problem that users find it difficult to quickly understand the content of the media file and improves the efficiency of information acquisition.

WO2025092132A1PCT designated stage expired Publication Date: 2025-05-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112721
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-08-16
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

After users share media files, view them to understand their content is inefficient, especially audio or video files, making it difficult for users to quickly understand the key information about the file.

Method used

Provide a data processing method. The server receives parsing instructions sent by the client, obtains the content information of the target media file, and sends the target information to the target model, receives and processes the output results of the target model, and finally sends display information to the client so that the user can quickly understand the content of the media file.

Benefits of technology

By automatically parsing and processing media files, users can quickly understand the content of the file, improve the efficiency of information acquisition, and reduce the need for file viewing frame by frame or syllable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112721_08052025_PF_FP_ABST
    Figure CN2024112721_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a data processing method and apparatus, and a storage medium, used for automatically analyzing a media file sent by an IM software client. Specifically, a server of IM software can receive an analysis instruction sent by a client for a target media file; upon acquisition of the analysis instruction, content information corresponding to the target media file can be acquired; then, target information generated according to the content information can be sent to a target model, such that the target model performs analysis on the basis of the target information, and returns an output result to the server; and upon acquiring the output result, the server can send display information to the client according to the output result. In this way, users can further understand the content of media files by means of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 31, 2023, with application number 202311433841.7 and invention name “A Data Processing Method, Device and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a data processing method, device, and storage medium. Background Art

[0003] In some scenarios, users share media files with other users. These media files can include audio files or video files. Alternatively, users can share media files by directly sending them to other users. Alternatively, users can share media files by sending the corresponding webpage to other users.

[0004] After obtaining a media file shared by other users, the user needs to view the media file to understand the content of the target media file, which is inefficient.

[0005] Summary of the Invention

[0006] In order to solve the problems of the prior art, the present application provides a data processing method, device and storage medium.

[0007] In a first aspect, the present application provides a data processing method, which is applied to a server and includes:

[0008] Receive parsing instructions for the target media file sent by the client;

[0009] Based on the parsing instruction, obtaining content information of the target media file;

[0010] Sending target information to a target model, where the target information is generated based on the content information;

[0011] receiving an output result returned by the target model based on the target information;

[0012] Sending display information to the client according to the output result.

[0013] In some possible implementations, obtaining content information corresponding to the target media file includes:

[0014] Acquire the target media file;

[0015] Extracting audio data of the target media file;

[0016] The audio data is converted into the content information.

[0017] In some possible implementations, obtaining content information corresponding to the target media file includes:

[0018] Obtaining a target link sent by the client, where the target link is a webpage link corresponding to the target media file;

[0019] Obtaining a subtitle file of the target media file according to the target link;

[0020] The content information is determined according to a subtitle file of the target media file.

[0021] In some possible implementations, the target information includes at least one prompt information;

[0022] Before sending the target information to the target model, the method further includes:

[0023] The at least one prompt information is generated according to the target information.

[0024] In some possible implementations, the at least one prompt information includes N segmented prompt information, where N is a positive integer greater than 1, and generating the at least one prompt information according to the target information includes:

[0025] Segmenting the content information to obtain N pieces of segmented content information;

[0026] Determine N pieces of segment prompt information according to the N segment content information;

[0027] The output result includes an output sub-result corresponding to each segment prompt information, and the output sub-result corresponding to the segment prompt information corresponds to the segment content information.

[0028] In some possible implementations, the segment prompt information further includes structured information corresponding to the N segment content information, where the structured information is used to indicate a structural relationship between the segment text content in the text content.

[0029] In some possible implementations, the at least one prompt information includes first prompt information, and generating the at least one prompt information according to the target information further includes:

[0030] First prompt information is generated based on the output sub-results generated by the N segmented prompt information.

[0031] In some possible implementations, the at least one prompt information includes second prompt information;

[0032] Generating at least one prompt information according to the content information includes:

[0033] extracting at least one target content information from the content information;

[0034] The second prompt information is generated based on the at least one target content information, the second prompt information is used to instruct the target model to extract information from the target content information, and the output result includes an output sub-result corresponding to the second prompt information, and the output sub-result corresponding to the second prompt information corresponds to the target content information.

[0035] In some possible implementations, the at least one prompt information includes third prompt information;

[0036] Generating at least one prompt information according to the content information includes:

[0037] Acquiring historical conversation content, where the historical conversation content includes information sent by the client and related to the target media file;

[0038] The third prompt information is generated according to the content information and the historical conversation content; the third prompt information is used to instruct the target model to obtain recommended content according to the historical conversation content and the content information, and the output result includes the recommended content.

[0039] In a second aspect, the present application provides a data processing method, which is applied to a client and includes:

[0040] Displaying a first content message, where the first content message is used to share a target media file or a link to the target media file;

[0041] Obtaining a parsing operation triggered for the target media file;

[0042] Send parsing instructions to the server;

[0043] Displaying the display information returned by the server, where the display information is determined according to the content of the target media file.

[0044] In some possible implementations, the parsing operation is triggered by any of the following:

[0045] An intelligent assistant invokes an operation on the target media file; or

[0046] A triggering operation is performed on a control displayed in association with the first content message.

[0047] In some possible implementations, the intelligent assistant invoking operation includes any one of the following:

[0048] An operation of sending the target media file to the intelligent assistant; or

[0049] An operation of sending an IM message for waking up an intelligent assistant in an instant messaging IM group chat session including the intelligent assistant, wherein the first content message is an IM message.

[0050] In some possible implementations, the displaying of the display information returned by the server includes:

[0051] generating a second content message according to the first content message, wherein the second content message includes information about the target media file and the presentation information;

[0052] The second content message is displayed at a display position corresponding to the first content message.

[0053] In some possible implementations, the presentation information includes summary information, where the summary information is obtained by summarizing the content of the target media file.

[0054] In some possible implementations, the presentation information includes recommended content; the recommended content is determined based on historical conversation content and the target media file, and the historical conversation content includes information about the target media file sent by the client.

[0055] In some possible implementations, the displaying of the display information returned by the server includes:

[0056] A recommended content control is displayed at an associated display position of the first content message, the recommended content control being used to display the recommended content and trigger a recommendation dialogue operation, the recommendation dialogue operation being used to send the recommended content to the server.

[0057] In some possible implementations, the method further includes:

[0058] In response to a forwarding operation triggered by the presentation information, the information of the target media file and the presentation information are forwarded.

[0059] In some possible implementations, the displaying of the display information returned by the server further includes:

[0060] At least one operation control is displayed, where the operation control is used to trigger a feedback operation, where the feedback operation is used to express the opinions of a user viewing the displayed information or to regenerate the displayed information.

[0061] In a third aspect, the present application provides a data processing device, wherein the method is applied to a server, including:

[0062] A first receiving unit, configured to receive a parsing instruction for a target media file sent by a client;

[0063] an acquiring unit, configured to acquire content information of the target media file based on the parsing instruction;

[0064] a first sending unit, configured to send target information to a target model, where the target information is generated based on the content information;

[0065] A second receiving unit is configured to receive an output result returned by the target model based on the target information;

[0066] The second sending unit is configured to send display information to the client according to the output result.

[0067] In some possible implementations, the acquisition unit is specifically configured to acquire the target media file; extract audio data from the target media file; and convert the audio data into the content information.

[0068] In some possible implementations, the acquisition unit is specifically configured to acquire a target link sent by the client, where the target link is a web page link corresponding to the target media file; acquire a subtitle file of the target media file based on the target link; and determine the content information based on the subtitle file of the target media file.

[0069] In some possible implementations, the target information includes at least one prompt information; the device further includes a prompt information generating unit; the prompt information generating unit is configured to generate the at least one prompt information according to the target information.

[0070] In some possible implementations, the at least one prompt information includes N segmented prompt information, where N is a positive integer greater than 1; the prompt information generation unit is specifically configured to segment the content information to obtain N segmented content information; determine N segmented prompt information based on the N segmented content information; and the output result includes an output sub-result corresponding to each segmented prompt information, and the output sub-result corresponding to the segmented prompt information corresponds to the segmented content information.

[0071] In some possible implementations, the segment prompt information further includes structured information corresponding to the N segment content information, where the structured information is used to indicate a structural relationship between the segment text content in the text content.

[0072] In some possible implementations, the at least one prompt information includes first prompt information, and the prompt information generating unit is specifically configured to generate the first prompt information based on the output sub-results generated by the N segmented prompt information.

[0073] In some possible implementations, the at least one prompt information includes second prompt information; the prompt information generation unit is specifically used to extract at least one target content information from the content information; the second prompt information is generated based on the at least one target content information, the second prompt information is used to instruct the target model to extract information from the target content information, and the output result includes an output sub-result corresponding to the second prompt information, and the output sub-result corresponding to the second prompt information corresponds to the target content information.

[0074] In some possible implementations, the at least one prompt information includes third prompt information; the prompt information generation unit is specifically used to obtain historical conversation content, and the historical conversation content includes information about the target media file sent by the client; the third prompt information is generated based on the content information and the historical conversation content; the third prompt information is used to instruct the target model to obtain recommended content based on the historical conversation content and the content information, and the output result includes the recommended content.

[0075] In a fourth aspect, the present application provides a data processing device, which is applied to a client and includes:

[0076] a display unit, configured to display a first content message, wherein the first content message is used to share a target media file or a link to the target media file;

[0077] an acquiring unit, configured to acquire a parsing operation triggered for the target media file;

[0078] A sending unit, used for sending parsing instructions to the server;

[0079] The display unit is further configured to display the display information returned by the server, where the display information is determined according to the content of the target media file.

[0080] In some possible implementations, the parsing operation is triggered by any one of the following: an intelligent assistant invoking operation for the target media file; or a triggering operation for a control displayed in association with the first content message.

[0081] In some possible implementations, the smart assistant invoking operation includes any one of the following: an operation of sending the target media file to the smart assistant; or an operation of sending an IM message for invoking the smart assistant in an instant messaging IM group chat session including the smart assistant, wherein the first content message is an IM message.

[0082] In some possible implementations, the display unit is specifically configured to generate a second content message based on the first content message, where the second content message includes information about the target media file and the display information; and display the second content message at a display position corresponding to the first content message.

[0083] In some possible implementations, the presentation information includes summary information, where the summary information is obtained by summarizing the content of the target media file.

[0084] In some possible implementations, the presentation information includes recommended content; the recommended content is determined based on historical conversation content and the target media file, and the historical conversation content includes information about the target media file sent by the client.

[0085] In some possible implementations, the display unit is further used to display a recommended content control at an associated display position of the first content message, the recommended content control is used to display the recommended content, and the recommended content control is further used to trigger a recommendation dialogue operation, and the recommendation dialogue operation is used to send the recommended content to the server.

[0086] In some possible implementations, the apparatus further includes a forwarding unit, configured to forward the information of the target media file and the presentation information in response to a forwarding operation triggered by the presentation information.

[0087] In some possible implementations, the display unit is further configured to display at least one operation control, where the operation control is configured to trigger a feedback operation, where the feedback operation is configured to express the views of a user viewing the display information or to regenerate the display information.

[0088] In a fifth aspect, the present application provides an electronic device, including:

[0089] one or more processors;

[0090] a storage device having one or more programs stored thereon,

[0091] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the methods described in the first aspect, or implement any of the methods described in the second aspect.

[0092] In a sixth aspect, the present application provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the program implements any of the methods described in the first aspect, or implements any of the methods described in the second aspect.

[0093] In a seventh aspect, the present application provides a computer program product, which, when running on a device, enables the device to execute the method described in the first aspect, or execute the method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0095] FIG1 is a flow chart of a data processing method provided in an embodiment of the present application;

[0096] FIG2 is another flow chart of a data processing method according to an embodiment of the present application;

[0097] FIG3 is a schematic diagram of an exemplary application scenario provided in an embodiment of the present application;

[0098] FIG4 is a schematic diagram of another exemplary application scenario provided by an embodiment of the present application

[0099] FIG5 is a schematic diagram of a data processing device provided in an embodiment of the present application;

[0100] FIG6 is a schematic diagram of another data processing device provided in an embodiment of the present application;

[0101] FIG7 is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0102] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0103] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0104] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0105] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different objects, devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0106] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0107] In some scenarios, users can share media files with other users. The user who has shared the target media file can view the media file shared by other users to understand the views or information that other users want to convey.

[0108] For example, when using IM software to communicate, users can share media files with other users through the IM software, and can also receive media files shared by other users through the IM client. For example, user A can send a video file or a voice IM message to user B through the IM session between them on the client used by user A. User B can then receive the video file or voice IM message sent by user A through the client used by user B. For another example, user A can send an IM message containing a link to user B through the client used by user A. User B can click the link in the IM message on the client to view the corresponding video file or audio file.

[0109] In the above process, user B needs to view the media file shared by user A to understand its content. However, viewing both audio and video files takes a considerable amount of time. Furthermore, due to the continuous nature of media files, fast-forwarding or skipping through parts of the file may result in missing key information.

[0110] In order to solve the problems of the prior art, an embodiment of the present application provides a method for processing data, which is described in detail below in conjunction with the drawings in the specification.

[0111] Referring to FIG. 1 , FIG. 1 is an information interaction diagram of a data processing method provided by an embodiment of the present application. This embodiment of the present application is applicable to application scenarios in which a user uses a server to parse media files. The data processing method can be executed by a data processing device. The data processing device can be implemented by software.

[0112] Optionally, the data processing device can be integrated into the server side of the software. The server side of the software can run on a server or a server cluster to provide corresponding services for the client side of the software. The client side of the software runs on a terminal device for use, such as a mobile phone or a computer. If the data processing device is integrated into the server side, the data processing device can communicate with the client side to obtain the parsing instructions sent by the client side, and send output results to the client side. Optionally, the data processing device can also communicate with the target model to send sample data and target request data to the target model. In an embodiment of the present application, the above-mentioned server side can be the server side of the IM software, and the above-mentioned client side can be the client side of the IM software. For the convenience of introduction, the following description will be made by taking the data processing device integrated into the server side of the IM software as an example.

[0113] As shown in FIG1 , the method specifically includes the following steps:

[0114] S101: Receive a parsing instruction for a target media file sent by a client.

[0115] To parse a target media file, the server can receive a parsing instruction for the target media file from the client. The parsing instruction is sent by the client using user instructions and represents the user's intention to parse the target media file. Based on the parsing instruction, the server can invoke the target model to parse the target media file.

[0116] In the embodiments of the present application, the target media file can be an audio file, a video file, or a link to a video file or an audio file. In other words, the server can obtain the target media file or the link to the target media file. For example, if user A shares a link to video C with user B, user B can trigger a parsing instruction for video C on the client used by user B.

[0117] Optionally, the parsing instruction may be obtained based on an operation triggered by the user on the client. Two possible implementations of triggering the parsing instruction are schematically introduced below.

[0118] In a first possible implementation, the parsing instruction can be obtained based on the action triggered by the parsing control. Specifically, the client can display the parsing control to the user on the interface. The user can trigger a click or other action on the parsing control. After obtaining the action triggered by the user on the parsing control, the client can generate the parsing instruction and send it to the server.

[0119] If the parsing instruction is obtained based on the parsing control, the parsing control can be displayed in association with the target media file, indicating that the parsing instruction for the target media file can be triggered by the parsing control. Optionally, the parsing control can be displayed in association with the target media file.

[0120] In some possible implementations, the parsing control can be displayed in a message card that carries the target media file. A message card is a type of message. The key information of the message can be organized into a message card according to a preset format and displayed to the user. Displaying the key information in the message card facilitates the user's access to the key information. Optionally, if the server is an IM software-level A server, the message card can be an IM message card.

[0121] For example, if the target media file is shared by user A to user B, a message card corresponding to the target media file may be displayed on the client of user B in the IM session corresponding to user A.

[0122] To invoke the target model to parse the target media file, a parsing control can be displayed in the message card. For example, the parsing control can be displayed at the bottom of the message card. Optionally, the message card can also include functional indication information. The functional indication information is used to indicate the purpose of the parsing control. User B can understand the purpose of the parsing control by reading the functional indication information, thereby knowing that the parsing control can be used to invoke the target model to parse the target media file and trigger the parsing control.

[0123] The message card corresponding to the target media file may include relevant information about the target media file. For example, it may include a preview of the target media file, a control for playing the target media file, etc. Optionally, if the target media file is shared as a link, that is, user A sends a link corresponding to the target media file to user B, the message card may also include the link corresponding to the target media file.

[0124] In a second possible implementation manner, the parsing instruction may be obtained according to the operation of sending a message.

[0125] Specifically, the parsing instruction can be obtained based on the user sending the target media file to the intelligent assistant. The intelligent assistant can be a pre-set object corresponding to the target model. The message sent by the user to the object can be parsed by the target model.

[0126] For an introduction to this part, please refer to the following, and I will not go into details here.

[0127] S102: Based on the parsing instruction, obtain content information of the target media file.

[0128] After receiving the parsing instruction, the server can first determine the content information of the target media file. The content information of the target media file corresponds to the audio data of the target media file. Alternatively, the content information of the target media file can be text data corresponding to the audio data of the target media file, and the text in the content information is consistent with the content described by the audio data of the target media file. The content information can be text information, which can also be referred to as text content information.

[0129] In the embodiment of the present application, the audio data of the target media file can be converted to obtain content information, and the content information of an existing target media file can also be obtained.

[0130] In a first possible implementation manner, the server may convert the audio data of the target media file to obtain content information of the target media file.

[0131] Specifically, the server may extract audio data from the target media file and convert the audio data into text to obtain content information of the target media file.

[0132] As described above, the target media file can be an audio file or a video file.

[0133] If the target media file is an audio file, then the target media file can be regarded as audio data. The server can parse the audio file to convert the audio file into content information.

[0134] If the target media file is a video file, and the target media file includes image data and audio data, the server can parse the target media file and extract the audio data from the target media file. Then, the server can convert the audio data into text to obtain the content information of the target media file.

[0135] As described above, the server can obtain the target media file sent by the client, and can also obtain the link to the target media file.

[0136] If the server obtains the target media file sent by the client, the server can store the target media file locally on the server. After obtaining the parsing instruction, the server can parse the locally stored target media file to obtain the content information of the target media file.

[0137] If the client sends a link to a target media file to the server, then after receiving the parsing instruction, the server can first retrieve the target media file based on the link, and then parse the audio data of the target media file to obtain the content information of the target media file. Optionally, the server corresponding to the link to the target media file can have a communication interface with the IM software server. The server can use this communication interface to send a retrieval request to the server corresponding to the link to the target media file, thereby retrieving the target media file.

[0138] In some possible implementations, the server may have the ability to convert audio data. Accordingly, after extracting the audio data of the target media file, the server may parse the audio data locally to obtain text content information corresponding to the target media file.

[0139] Alternatively, in some other possible implementations, the server may also convert the text content information by calling a service. For example, the server may be deployed with an interface corresponding to calling an audio-to-text cloud service, and the server may call the audio-to-text cloud service through the interface to convert the audio data of the target media file into text content information.

[0140] S103: Send target information to the target model.

[0141] After acquiring the content information, the server generates target information based on the content information and sends it to the target model. The target model parses the received target information to generate output results corresponding to the target information. Because the target information is determined based on the content information, the content information corresponds to the target media file, and accordingly, the output results also correspond to the target media file.

[0142] The target model may be a model with natural language processing (NLP) capabilities, such as a language model (LM). The target model is capable of parsing natural language input to determine the information corresponding to the natural language input. For example, the target model may summarize the natural language input to determine the core content of the natural language input. Alternatively, in some implementations, the target model may also perform other processing based on instructions, such as recommending other content related to the input based on a recommendation instruction.

[0143] In some possible implementations, the server may send target information to the server via prompt information. For example, the target information may include at least one prompt information. Specifically, the server may generate at least one prompt information based on the content information and send the at least one prompt information to the target model. The target model may process each prompt information separately, obtain a processing result corresponding to each prompt information, and return the output result to the server.

[0144] In the embodiments of the present application, the server can generate a variety of different prompt messages, each of which can have different functions. Depending on the input of different prompt messages, the processing results returned by the target model can have different functions. The following describes in detail five specific implementation methods as examples.

[0145] In the first implementation, the server can generate a full-text prompt message.

[0146] In a first possible implementation, the server can call the target model to summarize the full text of the content information. Accordingly, the server can generate a full-text prompt message based on the content information and send the full-text prompt message to the target model. The full-text prompt message includes the complete content information.

[0147] The target model can summarize the content information based on the full-text prompt information. Thus, the output structure returned by the target model includes a summary of the full-text content information. Based on the output of the target model, users can view the relevant information summarized from the complete target media file.

[0148] In actual scenarios, the amount of data that can be input into the target model may be limited. Accordingly, the amount of data that a prompt message can carry is also limited. If the length of the complete content information exceeds the upper limit of the amount of data that a single prompt message can carry, the server can divide the complete content information into multiple segments of content information, and then carry the multiple segments of content information in multiple full-text prompt messages and send them to the target model. The target model can summarize each full-text prompt message and return multiple intermediate results to the server. The server can then generate a new full-text prompt message based on the intermediate results and send the new prompt message to the target model. Based on the new prompt message, the server can summarize the intermediate results, thereby completing the summary of the full text of the content information. In this way, even if the length of the content information exceeds the length that a single prompt message can carry, the summary of the full text of the content information can be achieved.

[0149] In the second implementation, the server can generate segmented prompt information.

[0150] In some scenarios, the target media file may contain extensive content. Accordingly, the content information obtained from the target media file may also encompass multiple aspects. Therefore, if the complete content information is input into the target model, the target model may lose some or all aspects of the content information during the summarization process. In other words, directly summarizing the complete content information using the target model may result in the loss of relevant information in the target media file.

[0151] To this end, in some possible implementations, the server can segment the content information, generate multiple prompt messages based on the segmented content information, and then send these multiple prompt messages to the target model. This allows the target model to summarize each prompt message separately. Because the content information is segmented, the amount of text content carried in each prompt message is relatively small, and the target model will not lose information in the prompt during the summarization process.

[0152] In the embodiments of the present application, the information obtained by segmenting the content information can be referred to as segmented content information. The prompt information generated based on the segmented content information can be referred to as segmented prompt information. In other words, the content information can be divided into N segmented content information (N is a positive integer greater than 1), and then N segmented prompt information can be obtained based on the segmented content information.

[0153] Optionally, when segmenting content information, segmentation can be performed based on logical relationships between content information. For example, portions of content information describing identical or similar content can be grouped into separate content segments. Alternatively, content information can be segmented based on keywords. For example, portions containing the same keyword can be grouped into separate content segments, or portions where the keyword appears more frequently than a preset threshold can be grouped into separate content segments.

[0154] Optionally, the content information can be segmented with reference to the target media file. Because the content information is derived from the audio data of the target media file, the target media file segmentation method can be applied to the content information segmentation method. For example, pauses in the audio data of the target media file can be monitored, and the content information corresponding to the audio data before the pause and the content information after the pause can be divided into two different segmented content information.

[0155] Optionally, the content information may be segmented based on the audio features of the target media file. The audio features of the target media file may be characteristics of the audio data of the target media file, such as the timbre, frequency domain, and noise level of the audio data of the target media file. Alternatively, if the target media file includes voices of multiple people, the voices of different people may be segmented into different parts based on the timbre of the audio data of the target media file, with each part of the audio data corresponding to a segment of content information.

[0156] It should be noted that in the first implementation described above, the content information can also be split into multiple segments. However, while the implementation corresponding to the full-text prompt information divides the content information based on its length, the implementation corresponding to the segmented prompt information can divide the content information based on the actual situation of the target media file. In this way, the segmented prompt information, and even the structured prompt information described below, references the actual situation of the target media file, and the resulting output can more accurately describe the target media file.

[0157] If at least one prompt message generated by the server includes N segmented prompt messages, then after the target model obtains the N segmented prompt messages, it can send N segmented prompt messages to the target model to obtain the output sub-result corresponding to each prompt message.

[0158] In a third implementation, the server may generate structured prompt information. Optionally, the structured prompt information may also be referred to as first prompt information.

[0159] By segmenting the content information, users can view the key information corresponding to each audio segment in the target media file. However, this does not satisfy the user's need to view the complete key information of the target media file. To this end, the server can also generate structured prompt information so that the target model can determine the complete key information of the target media file based on the structured prompt information.

[0160] Specifically, the structured prompt information corresponds to the N segmented content information to ensure that the output result returned by the target model is obtained based on the complete content information. Optionally, the structured prompt information can include the N segmented content information.

[0161] Furthermore, the structured prompt information includes structured information. This structured information describes the relationships between the segmented content information. This allows the target model to determine the meaning and role of the segmented content information within the complete content information based on the relationships between the segmented content information, thereby determining the key information of the complete target media file.

[0162] Alternatively, the structured information may include, for example, sequence numbers of the segmented content information. The sequence numbers are used to indicate the order of the segmented content information in the content information. Alternatively, the structured information may also be used to describe the logical relationship between the segmented text contents.

[0163] Optionally, the structured prompt information may include complete content information, and structured information may be obtained by segmenting the content information. Accordingly, the target model may be better summarized based on the structured prompt information, and the obtained output structure may better reflect the key information of the complete target media file.

[0164] As mentioned above, a single prompt message may not be able to carry content information. Therefore, when sending a structured prompt message to a target model, if the content information is too large to carry N segments of content information within the structured prompt message, the server can send these N segments of content information to the target model via multiple sub-prompt messages. Optionally, the sub-prompt message can be a segmented prompt message. Alternatively, the sub-prompt message can include multiple segmented prompt messages.

[0165] Specifically, the server can split N segmented content information into M sub-prompt information (M is a positive integer greater than 1 and less than N). Each sub-prompt information includes at least one segmented content information and sub-structured information. The sub-structured content is used to describe the association relationship between the segmented content information included in the sub-prompt information. Different sub-prompt information includes different segmented content information.

[0166] That is, the i-th sub-prompt information among the M sub-prompt information includes j segments of content information and sub-structured information. The sub-structured information in the i-th sub-prompt information is used to indicate the association relationship between the j segments of text content. Where i is a positive integer less than or equal to M, and j is a positive integer less than N.

[0167] After obtaining M sub-prompt information, the server can send M sub-prompt information to the target model. The target model can obtain M sub-output results based on the M sub-prompt information and return them to the server. Each sub-output result corresponds to a sub-prompt information. The server can determine the structured prompt information based on the M sub-output results and the structured information. In this way, the server summarizes the sub-prompt information so that the length of the sub-output result is shorter than the length of the corresponding sub-prompt information. In this way, the segmented content information is compressed, so that the structured prompt information can carry the relevant content of multiple segmented content information.

[0168] In a fourth implementation, the server may generate a target extraction prompt message. Optionally, the target extraction prompt message may also serve as a second prompt message.

[0169] In some scenarios, there may be a need to summarize specific parts of the target media file. For example, if the target media file is used to introduce something, then the important content related to that thing in the target media file can be summarized. To this end, the server can generate a target extraction prompt message.

[0170] The target extraction prompt information is used to instruct the target model to summarize a specific part of the content information. Specifically, the target extraction prompt information includes at least one target content information. The target content information is a part of the content information that the target model needs to focus on summarizing. Optionally, if the server determines multiple target content information, the server can send the multiple target content information in one target extraction prompt information, or can send the multiple target content information in multiple target extraction prompt information.

[0171] Specifically, after obtaining the content information corresponding to the target media file, the server can extract the content information to extract the target content information from the content information. In some possible implementations, rules for extracting the target content information can be pre-configured. The server can extract the target content information according to the pre-configured rules.

[0172] In addition to the target content information, the target extraction prompt information may also include instruction information corresponding to the action of "target extraction" (hereinafter referred to as first instruction information). The first instruction information is used to instruct the target model to summarize the target content information.

[0173] In a fifth implementation, the server may generate a recommendation prompt message. Optionally, the recommendation prompt message may also be referred to as third prompt message.

[0174] In some scenarios, users may need recommendations. For example, after viewing the summary of a target media file, the user may have questions. To this end, the server can use the target model to predict the user's likely questions and present them to the user, eliminating the need for the user to re-enter the question. For another example, the user may also want to view media files similar to the target media file. Each time, the server can use the target model to predict the search keywords used by the user when searching for media files and present them to the user.

[0175] To this end, the server can also generate recommendation prompt information. This recommendation prompt information can include historical conversation content and second instruction information. The second instruction information is instruction information corresponding to the "recommend" action. The second instruction information is used to instruct the target model to make recommendations based on the target text content and historical conversation content. In some possible implementations, the second instruction information can include multiple instruction information. Different instruction information is used to instruct the target model to recommend different content.

[0176] The historical conversation content can be conversations between the user and the intelligent assistant. By analyzing the historical conversation content, the target model can determine the user's intent and make recommendations based on the user's intent. Alternatively, the historical conversation content can also be conversations between users.

[0177] It is understood that in actual application scenarios, the server can generate any one or more of the five prompt messages mentioned above. In addition, the server can also generate other types of prompt messages. We will not elaborate on them here.

[0178] S104: Receive the output result returned by the target model based on the target information.

[0179] After receiving the target information sent by the server, the target model can process the target information to obtain the corresponding output result and return it to the server. The server can receive the output result returned by the target model based on the target information.

[0180] As previously described, the server can generate multiple prompt messages based on the content information and send them to the target model. Accordingly, the server's output result can include multiple output sub-results. Different output sub-results can correspond to different prompt messages. In other words, based on different prompt messages, the target model can return different output sub-results.

[0181] The above describes five possible prompt messages. The following describes the corresponding output sub-results for each of these five prompt messages. It is understood that if the server also generates other types of prompt messages, the target model can return the corresponding output sub-results.

[0182] In the first implementation, the server sends a full-text prompt message to the target model, and the output sub-results include the output sub-results corresponding to the full-text prompt message. Accordingly, the target model can summarize the content information based on the full-text prompt message. The output sub-results corresponding to the full-text prompt message can include a summary of the complete content information.

[0183] If the server sends the full-text prompt information to the target model, the output result returned by the target model may include the output sub-result corresponding to the full-text prompt information, or the output result returned by the target model may be further generated based on the output sub-result corresponding to the full-text prompt information.

[0184] In a second implementation, the server can send N segmented prompt messages to the target model, each of which can contain a portion of the content information. An output result can then be generated based on the output sub-results corresponding to each of the N segmented prompt messages. Specifically, the target model can summarize each segmented prompt message to obtain a summary corresponding to each segmented content information. It is understood that the output sub-result corresponding to a segmented prompt can include a summary corresponding to the segmented content information in the segmented prompt message.

[0185] If the server sends N segmented prompt information to the target model, the output result returned by the target model may include the output sub-result corresponding to each segmented prompt information, or the output result returned by the target model may be further generated based on the output sub-result corresponding to the segmented prompt information.

[0186] In a third implementation, the server can send structured prompt information to the target model, and the output result includes an output sub-result corresponding to the structured prompt information. The target model can summarize the multiple segmented content information based on the structured information in the structured prompt information. In this way, the output sub-result corresponding to the structured prompt information includes a summary of the complete content information that references the structure of the content information.

[0187] If the server sends structured prompt information to the target model, the output result returned by the target model may include the output sub-result corresponding to the structured prompt information, or the output result returned by the target model may be further generated based on the output sub-result corresponding to the structured prompt information.

[0188] In a fourth implementation, the server sends a target extraction prompt message to the target model, and the output result includes an output sub-result corresponding to the target extraction prompt message. The target model can summarize the target content information in the target extraction prompt message based on the first instruction message. In this way, the output sub-result corresponding to the target extraction prompt message includes a summary of the target text content.

[0189] If the server sends target extraction prompt information to the target model, the output result returned by the target model may include the output sub-result corresponding to the target extraction prompt information, or the output result returned by the target model may be further generated based on the output sub-result corresponding to the target extraction prompt information.

[0190] In the fifth implementation, the server sends a recommendation prompt to the target model. The output includes a sub-result corresponding to the recommendation prompt. The target model can learn the context of the historical conversations contained in the recommendation prompt and, after learning the context, determine the recommended content to recommend to the user. The sub-result corresponding to the recommendation prompt includes the recommended content.

[0191] If the server sends recommendation prompt information to the target model, the output result returned by the target model may include the output sub-result corresponding to the recommendation prompt information, or the output result returned by the target model may be further generated based on the output sub-result corresponding to the recommendation prompt information.

[0192] The above describes five types of output sub-results. In practical scenarios, the target model can generate corresponding output sub-results based on the prompt information sent by the server. If the server generates multiple prompt information based on a media file, the target model can generate multiple output sub-results based on these multiple prompt information. These multiple output sub-results can be collectively referred to as output results. In other words, the output sub-results generated based on the target media file can be collectively referred to as the output result corresponding to the target media file.

[0193] Alternatively, the server can further process the multiple output sub-results to obtain the final output result. For example, the server can summarize the multiple output sub-results or re-call the target model to summarize the multiple output sub-results.

[0194] S105: Sending display information to the client according to the output result.

[0195] After obtaining the output results returned by the target model, the server can generate presentation information based on the output results and send the presentation information to the client, which then displays the presentation information to the user. Optionally, the presentation information can include summary information. The summary information summarizes the content of the target media file.

[0196] Optionally, the client may display the display information to the user at a display location associated with the IM message corresponding to the target media file. For example, if the client displays the target media file to the user via an IM message card, the client may demarcate a display area in the IM message card for displaying the output result, and display the output result in the display area.

[0197] Furthermore, the server can associate and store the presentation information with the target media file. This way, if another client triggers a command to parse the target media file, the server can return the stored presentation information. Alternatively, if the client forwards the target media file to another user, the server can simultaneously forward the target media file and the presentation information.

[0198] Specifically, the server can associate the output result with the IM message card corresponding to the target media file. In this way, when the target media file is forwarded to other clients, the server can send the IM message card to the corresponding client to display the target media file and display information to the user.

[0199] In one implementation, the server may receive user input associated with a target media file from the client and send the user input to the target model. The target model generates second presentation information based on the content information of the target media file and the user input, and sends the second presentation information to the server. The server receives the second presentation information returned by the target model and sends the second presentation information to the client.

[0200] The present application provides a data processing method, device, and storage medium for automatically parsing media files sent by an IM software client. Specifically, the server of the IM software can receive a parsing instruction for a target media file sent by the client. After obtaining the parsing instruction, the content information corresponding to the target media file can be obtained. Then, the target information generated based on the content information can be sent to the target model so that the target model can perform parsing based on the target information and return the output result to the server. After obtaining the output result, the server can send display information to the client based on the output result. In this way, if the IM client receives relevant information about the target media file, the user can trigger the parsing instruction on the IM client. According to the parsing instruction triggered by the user, the server can first obtain the content information corresponding to the target media file, and then send the content information to the target model for parsing. In this way, the user can use the model to further understand the content of the media file.

[0201] The above describes some implementation methods for parsing target media files by calling the target model from the server's perspective. The following describes how to trigger the server to parse the target media file from the client's perspective.

[0202] See Figure 2, which is a flow chart illustrating a data processing method provided herein. This embodiment of the present application is applicable to scenarios in which a user uses a client to invoke a server to parse a media file. The data processing method can be executed by a data processing device. The data processing device can be implemented by software running on a server.

[0203] Optionally, the data processing device may be integrated into a software client. The software client may run on a terminal device to provide corresponding services to users. In the embodiments of the present application, the aforementioned server may be an IM software client. For ease of description, the following description uses the example of a data processing device integrated into an IM software client.

[0204] As shown in FIG2 , the method specifically includes the following steps:

[0205] S201: Display a first content message.

[0206] In an embodiment of the present application, the target media file to be parsed comes from a client. Specifically, before parsing the target media file, the client may first display a first content message. The first content message is used to share the target media file or a link to the target media file. The link to the target media file may be a URL corresponding to the target media file, such as a Uniform Resource Locator (URL) corresponding to the target media file.

[0207] Optionally, the first content message can be a message sent by a user using the client through the client, or a message sent by another user to the user using the client. That is, a user can send a target media file or share a link to a target media file to another user through the client. Alternatively, a user can receive a target media file sent by another user or receive a link to a target media file shared by another user through the client. Optionally, the aforementioned other users can include an intelligent assistant. For an introduction to intelligent assistants, please refer to the above and will not be repeated here.

[0208] In some possible implementations, the client is an IM client, and the first content message may be an IM message. Accordingly, the client may display the target media file via an IM message card. The first content message may be an IM message card used to display the target media file. Optionally, the IM message card may display information such as the cover and title of the target media file. If the first content message is used to share a link to the target media file, the IM message card may also include the link to the target media file. If the first content message is used to share the target media file, the IM message card may also include a download control for triggering the download of the target media file.

[0209] S202: Obtaining a parsing operation triggered for a target media file.

[0210] After the first content message is displayed to the user, the user may trigger a parsing operation for the target media file. According to the parsing operation triggered by the user, the client may call the target model through the server to parse the target media file.

[0211] In the embodiment of the present application, the user can trigger the parsing operation in a variety of ways. The following describes some ways to trigger the parsing operation.

[0212] In a first possible implementation, the parsing operation may be triggered by an intelligent assistant invoking operation on the target media file.

[0213] Specifically, the user can invoke the intelligent assistant so that the target media file can be parsed by the intelligent assistant. Alternatively, the user can invoke the intelligent assistant by sending a message to the intelligent assistant. That is, the user's operation of invoking the intelligent assistant may include the user's operation of sending a message to the intelligent assistant.

[0214] Optionally, the user sending a message to the intelligent assistant may include the user sending an IM message to invoke the intelligent assistant. Specifically, the user may send an IM message to invoke the intelligent assistant in an IM group chat session that includes the intelligent assistant. The IM message may include an invocation instruction and the name of the intelligent assistant in the IM group chat. The invocation instruction may be "@". Accordingly, the first content message may be a message sent in the IM group chat session.

[0215] That is to say, for media files or links to media files shared in an IM group chat session, if the IM group chat session includes an intelligent assistant, the user can send an IM message in the IM group chat session to invoke the intelligent assistant, so that the media file can be parsed by the intelligent assistant.

[0216] Alternatively, the user sending a message to the intelligent assistant may also include the user sending a message in an IM single chat session with the intelligent assistant. In other words, the user sending the target media file to the intelligent assistant is equivalent to the intelligent assistant invoking the target media file. Accordingly, the above-mentioned first content message may be a message sent by the user to the intelligent assistant. After obtaining the target media file or the link to the target media file that the user has asked the intelligent assistant to analyze, the client may display the first content message and obtain the parsing instruction for the target media file.

[0217] In a second implementation manner, the user may trigger a parsing operation through a space displayed in association with the first content message.

[0218] Specifically, the client may display an operation control in association with the first content message. The user may trigger the parsing operation by triggering the operation control. For example, assuming the first content message includes an IM message card, the IM message card may include an operation control for triggering the parsing operation, so that the user may trigger the parsing operation by triggering the operation control.

[0219] 3 , which is a schematic diagram of an application scenario provided by an embodiment of the present application. For example, when a user uses a client installed on a terminal device (such as a mobile phone or a computer), the client may display a display area 300 as shown in FIG3 .

[0220] Specifically, the display area 300 is used to display relevant content in the IM group chat. In the implementation shown in FIG3 , the IM group chat is named “Learning Materials Exchange Group” and the user using the client is user B.

[0221] In the implementation shown in Figure 3, display area 300 is used for IM messages 310 and 320. IM message 310 is sent by user A in an IM group chat, sharing a video with other users in the group chat. IM message 320 is sent by user B in the group chat, initiating the intelligent assistant to analyze the video shared by user A. In the implementation corresponding to Figure 3, the title of the video shared by user A is "Learning Video - Explanation and Analysis of AAA Knowledge Points."

[0222] In order to indicate the user who sends the IM message, in the implementation shown in FIG. 3 , the IM message may be displayed in association with the name and avatar of the user who sends the IM message.

[0223] Group chat message 310 is an IM message card that displays information about the video shared by user A. Specifically, group chat message 310 includes a link 311 to the target media file and a cover 312 for the target media file. Other users in the group chat can view the cover of the target media file through group chat message 312 and can also click link 311 to jump to the corresponding page of the video.

[0224] In addition, group chat message 310 also includes a parsing control 313. If group chat message 310 is regarded as the "first content message" mentioned above, parsing control 313 can be equivalent to the "control displayed in association with the first content message" mentioned above. The user can trigger the parsing operation by triggering parsing control 313.

[0225] In the implementation shown in Figure 3, user B can trigger the operation of invoking the intelligent assistant by sending an IM message 320, thereby invoking the target model to parse the video shared by user A. That is, if group chat message 310 is regarded as the "first content message" mentioned above, IM message 320 can be equivalent to the "IM message sent in an IM group chat session including the intelligent assistant to invoke the intelligent assistant" mentioned above.

[0226] S203: Send parsing instructions to the server.

[0227] After receiving the parsing instructions, the client can send them to the server. The server can determine the content of the target media file based on the parsing, generate target information based on the content information, and send it to the target model. The server then receives the output from the target information, determines presentation information based on the output, and returns the presentation information to the server. For an introduction to this aspect, please refer to the implementation shown in Figure 1 and will not be further elaborated here.

[0228] S204: Display information returned by the display server.

[0229] After receiving the presentation information returned by the server, the client can display it. The presentation message can include summary information. This summary information summarizes the contents of the target media file. This allows the user to understand the general content of the target media file by viewing the presentation message. Optionally, the presentation information can also include other information. For details on this, please refer to the description of the embodiment corresponding to FIG1 and will not be repeated here.

[0230] In an embodiment of the present application, the parsed target media file may be shared via a first content message. Therefore, when displaying the presentation information, the presentation information may be displayed via the content message. Specifically, the client may generate a second content message based on the first content message and the presentation information, and display the second content message at the display location corresponding to the first content message, thereby displaying the presentation content via the second content message. For example, if the first content message is an IM message, the client may generate an IM message including the presentation content based on the first content message, and display the IM message including the presentation content at the location corresponding to the first content message.

[0231] The implementation shown in Figure 3 is used as an example for description. After user B sends IM message 320, the client can obtain the display information and display the interface shown in Figure 4. In the implementation shown in Figure 4, the client can display the display area 400 shown in Figure 4.

[0232] Specifically, display area 400 includes IM message 410. IM message 410 is obtained based on the display information and IM message 310 in the implementation shown in FIG3. Moreover, the display position of IM message 410 matches IM message 310, which is equivalent to replacing IM message 310 with IM message 410.

[0233] Specifically, IM message 410 includes a link 411 to the target media file, a target media file cover 412, and presentation information 413. Presentation information 413 is derived based on the content information of the target media file. In the implementation corresponding to FIG4 , presentation information 413 includes summary information such as "This video explains..." By viewing presentation information 413, the user can understand the general content of the target media file.

[0234] In some possible implementations, a user can trigger a forwarding operation based on the presentation information. For example, a user can trigger a forwarding operation based on IM message 410 in the implementation shown in FIG. 4 . After receiving the user-triggered forwarding operation, the target media file information and presentation information can be forwarded. Specifically, IM message 410 can be forwarded in its entirety. That is, IM message 410 can be copied to another IM session.

[0235] As previously described, the prompt information sent by the server to the target model may include recommendation prompt information. The output sub-result returned by the target model includes recommended content for the user. If the output result includes recommended content, the presentation information may also include the recommended content. Accordingly, when displaying the presentation information, the client may display a recommended content control at the associated display location of the first content message. The recommended content control is used to display recommended content to the user.

[0236] Furthermore, the recommended content control is also used to trigger a recommended conversation operation. After receiving the triggered recommended conversation operation, the client can send recommended content to the server. In other words, the recommended content can be questions that the target model predicts the user might want to ask after viewing the summary information (or other information). The client can present these questions to the user in the form of a recommended content control. The user can trigger the recommended content control to ask the target model another question. This way, the user's likely questions can be predicted, eliminating the need for the user to manually enter the question.

[0237] For example, in the implementation corresponding to FIG4 , the display area 400 may further include a recommended content control 420. The recommended content control 420 is used to display recommended content obtained by the target model. Specifically, the recommended content may include three questions that the user may ask: Question A, Question B, and Question C. The user can ask Question A to the target model by clicking the display area corresponding to Question A in the recommended content control 420; can also ask Question B to the target model by clicking the display area corresponding to Question B in the recommended content control 420; and can also ask Question C to the target model by clicking the display area corresponding to Question C in the recommended content control 420.

[0238] In some possible implementations, the user may be dissatisfied with the output of the target model. In this case, the user may need to express their opinion and / or instruct the target media file's content information to be reprocessed. In other possible implementations, the user may be very satisfied with the output of the target model. In this case, the user may also need a channel to provide feedback on their opinion.

[0239] To this end, in an embodiment of the present application, the client may further display at least one operation control. The user may trigger a feedback operation through the operation control. The feedback operation may express the user's opinion regarding the displayed information or instruct the server to regenerate the displayed information. Thus, by triggering the corresponding operation control, the user may provide feedback to the server regarding their opinion or instruct the server to regenerate the displayed information.

[0240] For example, in the implementation corresponding to FIG4 , IM message 410 also includes an operation control 414 and an operation control 415. Operation control 414 is used to trigger a re-parsing operation. A user can trigger operation control 414 to instruct the re-parsing of the content information of the target media file. Operation control 415 is used for the user to provide feedback on their own opinions. Specifically, operation control 415 includes two sub-controls. One sub-control is used to indicate that the user's feedback on the displayed content is satisfactory, and the other sub-control is used to indicate that the user is dissatisfied with the displayed content.

[0241] In one implementation, the user terminal may send user input associated with a target media file to the server terminal. In response to receiving the user input, the target model generates second presentation information based on the content information of the target media file and the user input. The server terminal sends the second presentation information returned by the target model to the client terminal. The client terminal receives the second presentation information and displays it.

[0242] Based on a data processing method provided in the above method embodiment, the embodiment of the present application also provides a data processing device applied to the server and a data processing device applied to the client. The data processing device will be described below in conjunction with the accompanying drawings.

[0243] Refer to Figure 5, which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The device is applied to the server, as shown in Figure 5, and the data processing device includes:

[0244] The first receiving unit 510 is configured to receive a parsing instruction for a target media file sent by a client;

[0245] An acquiring unit 520, configured to acquire content information of the target media file based on the parsing instruction;

[0246] A first sending unit 530 is configured to send target information to a target model, where the target information is generated based on the content information;

[0247] A second receiving unit 540 is configured to receive an output result returned by the target model based on the target information;

[0248] The second sending unit 550 is configured to send presentation information to the client according to the output result.

[0249] In some possible implementations, the acquisition unit 520 is specifically configured to acquire the target media file; extract audio data from the target media file; and convert the audio data into the content information.

[0250] In some possible implementations, the acquisition unit 520 is specifically configured to acquire a target link sent by the client, where the target link is a web link corresponding to the target media file; acquire a subtitle file of the target media file based on the target link; and determine the content information based on the subtitle file of the target media file.

[0251] In some possible implementations, the target information includes at least one prompt information; the device further includes a prompt information generating unit; the prompt information generating unit is configured to generate the at least one prompt information according to the target information.

[0252] In some possible implementations, the at least one prompt information includes N segmented prompt information, where N is a positive integer greater than 1; the prompt information generation unit is specifically configured to segment the content information to obtain N segmented content information; determine N segmented prompt information based on the N segmented content information; and the output result includes an output sub-result corresponding to each segmented prompt information, and the output sub-result corresponding to the segmented prompt information corresponds to the segmented content information.

[0253] In some possible implementations, the segment prompt information further includes structured information corresponding to the N segment content information, where the structured information is used to indicate a structural relationship between the segment text content in the text content.

[0254] In some possible implementations, the at least one prompt information includes first prompt information, and the prompt information generating unit is specifically configured to generate the first prompt information based on the output sub-results generated by the N segmented prompt information.

[0255] In some possible implementations, the at least one prompt information includes second prompt information; the prompt information generation unit is specifically used to extract at least one target content information from the content information; the second prompt information is generated based on the at least one target content information, the second prompt information is used to instruct the target model to extract information from the target content information, and the output result includes an output sub-result corresponding to the second prompt information, and the output sub-result corresponding to the second prompt information corresponds to the target content information.

[0256] In some possible implementations, the at least one prompt information includes third prompt information; the prompt information generation unit is specifically used to obtain historical conversation content, and the historical conversation content includes information about the target media file sent by the client; the third prompt information is generated based on the content information and the historical conversation content; the third prompt information is used to instruct the target model to obtain recommended content based on the historical conversation content and the content information, and the output result includes the recommended content.

[0257] See FIG6 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The device is applied to a client, as shown in FIG6 , and the data processing device includes:

[0258] A display unit 610 is configured to display a first content message, where the first content message is used to share a target media file or a link to the target media file;

[0259] An acquiring unit 620 is configured to acquire a parsing operation triggered for the target media file;

[0260] The sending unit 630 is used to send a parsing instruction to the server;

[0261] The display unit 610 is further configured to display the display information returned by the server, where the display information is determined according to the content of the target media file.

[0262] In some possible implementations, the parsing operation is triggered by any one of the following: an intelligent assistant invoking operation for the target media file; or a triggering operation for a control displayed in association with the first content message.

[0263] In some possible implementations, the smart assistant invoking operation includes any one of the following: an operation of sending the target media file to the smart assistant; or an operation of sending an IM message for invoking the smart assistant in an instant messaging IM group chat session including the smart assistant, wherein the first content message is an IM message.

[0264] In some possible implementations, the display unit 610 is specifically configured to generate a second content message based on the first content message, where the second content message includes information about the target media file and the display information; and display the second content message at a display position corresponding to the first content message.

[0265] In some possible implementations, the presentation information includes summary information, where the summary information is obtained by summarizing the content of the target media file.

[0266] In some possible implementations, the presentation information includes recommended content; the recommended content is determined based on historical conversation content and the target media file, and the historical conversation content includes information about the target media file sent by the client.

[0267] In some possible implementations, the display unit 310 is further used to display a recommended content control at an associated display position of the first content message, the recommended content control is used to display the recommended content, and the recommended content control is further used to trigger a recommendation dialogue operation, and the recommendation dialogue operation is used to send the recommended content to the server.

[0268] In some possible implementations, the apparatus further includes a forwarding unit, configured to forward the information of the target media file and the presentation information in response to a forwarding operation triggered by the presentation information.

[0269] In some possible implementations, the display unit 610 is further configured to display at least one operation control, where the operation control is configured to trigger a feedback operation, where the feedback operation is configured to express the views of a user viewing the display information or to regenerate the display information.

[0270] Based on a data processing method provided in the above method embodiment, the present application also provides an electronic device, including: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method described in any of the above embodiments.

[0271] Reference is now made to FIG7 , which shows a schematic diagram of the structure of an electronic device 700 suitable for implementing an embodiment of the present application. The terminal devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable Android devices), PMPs (Portable Media Players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs (television sets), desktop computers, and the like. The electronic device shown in FIG7 is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0272] As shown in Figure 7, electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of electronic device 700 are also stored in RAM 703. Processing device 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0273] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0274] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present application are performed.

[0275] The electronic device provided in the embodiment of the present application and the data processing method provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0276] Based on a data processing method provided in the above method embodiment, an embodiment of the present application provides a computer storage medium on which a computer program is stored, wherein when the program is executed by a processor, the data processing method described in any of the above embodiments is implemented.

[0277] It should be noted that the computer-readable medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0278] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0279] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0280] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the data processing method.

[0281] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0282] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0283] The units involved in the embodiments described in this application may be implemented by software or hardware. In some cases, the name of a unit / module does not constitute a limitation of the unit itself. For example, a voice data acquisition module may also be described as a "data acquisition module."

[0284] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0285] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0286] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0287] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0288] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0289] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, the method being applied to a server, comprising: Receive a parsing instruction for a target media file sent by a client; Based on the parsing instruction, obtaining content information of the target media file; Sending target information to a target model, wherein the target information is generated based on the content information; Receiving an output result returned by the target model based on the target information; Sending display information to the client according to the output result.

2. The method according to claim 1, wherein obtaining the content information corresponding to the target media file comprises: Acquire the target media file; Extracting audio data of the target media file; The audio data is converted into the content information.

3. The method according to claim 1, wherein obtaining content information corresponding to the target media file comprises: Acquire a target link sent by the client, where the target link is a webpage link corresponding to the target media file; According to the target link, obtaining a subtitle file of the target media file; The content information is determined according to the subtitle file of the target media file.

4. The method according to any one of claims 1 to 3, wherein the target information includes at least one prompt information; Before sending the target information to the target model, the method further includes: The at least one prompt information is generated according to the target information.

5. The method according to claim 4, wherein the at least one prompt information comprises N segmented prompt information, wherein N is a positive integer greater than 1, and wherein generating the at least one prompt information according to the target information comprises: Segmenting the content information to obtain N segmented content information; Determine N pieces of segment prompt information according to the N pieces of segment content information; The output result includes the output sub-results corresponding to each of the segmented prompt information. The output sub-result corresponding to the segment prompt information corresponds to the segment content information.

6. The method according to claim 5, wherein the segmentation prompt information further includes structured information corresponding to the N segmentation content information, and the structured information is used to indicate the structural relationship of the segmentation text content in the text content.

7. The method according to claim 6, wherein the at least one prompt information comprises first prompt information, and the generating the at least one prompt information according to the target information further comprises: Generate first prompt information based on the output sub-results generated by the N segmented prompt information.

8. The method according to claim 4, wherein the at least one prompt information comprises a second prompt information; Generating at least one prompt information according to the content information comprises: extracting at least one target content information from the content information; The second prompt information is generated according to the at least one target content information, the second prompt information is used to instruct the target model to extract information from the target content information, the output result includes an output sub-result corresponding to the second prompt information, and the output sub-result corresponding to the second prompt information corresponds to the target content information.

9. The method according to claim 4, wherein the at least one prompt information comprises third prompt information; Generating at least one prompt information according to the content information comprises: Acquire historical conversation content, where the historical conversation content includes information about the target media file sent by the client; Generate the third prompt information according to the content information and the historical conversation content; The third prompt information is used to instruct the target model to obtain recommended content based on the historical conversation content and the content information, and the output result includes the recommended content.

10. The method according to claim 1, wherein the method further comprises: receiving user input associated with the target media file sent by the client; Second presentation information is sent to the client, where the second presentation information is generated by the target model based on content information of the target media file and the user input.

11. A data processing method, the method being applied to a client, comprising: Display a first content message, the first content message is used to share the target media file or the target Links to media files; Obtaining a parsing operation triggered for the target media file; Send parsing instructions to the server; The display information returned by the server is displayed, where the display information is determined according to the content of the target media file.

12. The method according to claim 11, wherein the parsing operation is triggered by any one of the following: an intelligent assistant invoking an operation on the target media file; or A triggering operation on a control displayed in association with the first content message.

13. The method according to claim 12, wherein the intelligent assistant invoking operation comprises any one of the following: An operation of sending the target media file to the intelligent assistant; or An operation of sending an IM message for waking up an intelligent assistant in an instant messaging IM group chat session including the intelligent assistant, wherein the first content message is an IM message.

14. The method according to claim 11, wherein the displaying the display information returned by the server comprises: generating a second content message according to the first content message, wherein the second content message includes information of the target media file and the presentation information; The second content message is displayed at a display position corresponding to the first content message. 15 . The method according to claim 11 , wherein the presentation information comprises summary information, and the summary information is obtained by summarizing the content of the target media file.

16. The method according to claim 11, wherein the display information includes recommended content; the recommended content is determined according to historical conversation content and the target media file, and the historical conversation content includes information about the target media file sent by the client.

17. The method according to claim 16, wherein the displaying of the display information returned by the server comprises: A recommended content control is displayed at an associated display position of the first content message, the recommended content control is used to display the recommended content, and the recommended content control is also used to trigger a recommended dialogue operation, and the recommended dialogue operation is used to send the recommended content to the server.

18. The method according to claim 11, wherein the method further comprises: In response to a forwarding operation triggered by the presentation information, the information of the target media file and the presentation information are forwarded.

19. The method according to claim 14, wherein the displaying the display information returned by the server further comprises: At least one operation control is displayed, where the operation control is used to trigger a feedback operation, where the feedback operation is used to express the viewpoint of a user viewing the displayed information or to regenerate the displayed information.

20. The method according to claim 11, wherein the method further comprises: receiving user input associated with the target media file; Sending the user input to the server; The second display information returned by the server is displayed, where the second display information is generated based on the content of the target media file and the user input.

21. A data processing device, applied to a server, comprising: A first receiving unit, configured to receive a parsing instruction for a target media file sent by a client; An acquisition unit, configured to acquire content information of the target media file based on the parsing instruction; A first sending unit, configured to send target information to a target model, wherein the target information is generated based on the content information; A second receiving unit, configured to receive an output result returned by the target model based on the target information; The second sending unit is used to send display information to the client according to the output result.

22. A data processing device, applied to a client, comprising: A display unit, used to display a first content message, where the first content message is used to share a target media file or a link to the target media file; An acquisition unit, configured to acquire a parsing operation triggered for the target media file; A sending unit, used for sending a parsing instruction to a server; The display unit is further used to display the display information returned by the server, and the display information is determined according to the content of the target media file.

23. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 20.

24. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 20 is implemented.

Citation Information

Patent Citations

  • Video processing method and device, storage medium and electronic equipment

    CN110868632A

  • Video content recognition method and device, storage medium and electronic equipment

    CN113569610A

  • Multimedia content analysis method and device, equipment and storage medium

    CN116702749A

  • Video automatic analysis method and device based on large language model control and medium

    CN116935288A

  • Multimedia content sharing method and apparatus, device, and medium

    US20230300429A1