Communication file association analysis information generation method, device, equipment and readable medium
By collecting user-identified communication document association behavior profile information, personalized recommendation association types and file sets are generated, solving the problem of poor user experience in existing technologies and improving the efficiency and accuracy of users in cross-file content association and extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING CAIDODUI INFORMATION TECH CO LTD
- Filing Date
- 2025-08-04
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, when users associate and extract content across files, they rely on general rules or standards to select and associate files, failing to provide personalized recommendations based on users' historical behavior or preferences, resulting in a poor user experience.
By collecting user identification information related to communication file association behavior profiles, the system generates recommended association type information and file sets, provides personalized association options, and generates communication file association parsing information after the user makes a selection.
It improves the user experience by reducing the time and effort users spend filtering and associating files through personalized recommendations, and provides rich reference information to help users gain a more comprehensive understanding of file content and association possibilities.
Smart Images

Figure CN120994618B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of computer technology, and more specifically to a method, apparatus, device, and readable medium for generating communication document association parsing information. Background Technology
[0002] Communication document association resolution information generation is a technology that performs cross-file content association and extraction on multiple communication documents to generate communication document association resolution information. Currently, when performing cross-file content association and extraction on multiple communication documents, the common approach is as follows: users view various communication documents, select the communication documents that need to be associated, and then perform cross-file content association and extraction on the selected communication documents based on common rules or standards (such as file type, creation time, etc.) to obtain communication document association resolution information.
[0003] However, when using the above method to perform cross-file content association and extraction on multiple communication documents, the following technical problems often arise:
[0004] Users view various communication files, select the communication files they want to associate, and then perform cross-file content association and extraction based on common rules or standards (such as file type, creation time, etc.) to obtain communication file association parsing information. This process relies entirely on users to select and associate files themselves, without providing personalized recommendations based on users' historical behavior or preferences. This may cause users to spend more time and effort to filter and associate files, resulting in a poor user experience.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure provide a method, apparatus, electronic device, and computer-readable medium for generating communication document association parsing information to solve one or more of the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide a method for generating communication file association resolution information. The method includes: responding to receiving communication file association resolution task information sent by a client; collecting each communication file to be associated corresponding to the communication file association resolution task information, wherein the communication file association resolution task information includes a user identifier, and each communication file to be associated corresponds to a service object communicator identifier and a file identifier; obtaining communication file association behavior profile information corresponding to the user identifier; generating recommended association type information based on the communication file association behavior profile information; and generating a set of recommended associated communication files corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and the communication files to be associated. For each of the aforementioned communication files to be associated, generate corresponding communication file feature information, and determine the communication files to be associated and their feature information as display file information; determine the obtained display file information as a display file information set; send the recommended association type information, the recommended association file sets, and the display file information set to the client for the client to display; in response to receiving the filtering association interaction information sent by the client, generate communication file association parsing information based on the filtering association interaction information and the communication files to be associated; send the communication file association parsing information to the client.
[0009] Secondly, some embodiments of this disclosure provide a communication file association resolution information generation apparatus. The apparatus includes: a collection unit configured to, in response to receiving communication file association resolution task information sent by a client, collect each communication file to be associated corresponding to the communication file association resolution task information, wherein the communication file association resolution task information includes a user identifier, and each of the communication files to be associated corresponds to a service object communicator identifier and a file identifier; an acquisition unit configured to acquire communication file association behavior profile information corresponding to the user identifier; a first generation unit configured to generate recommended association type information based on the communication file association behavior profile information; a second generation unit configured to generate sets of recommended associated communication files corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and the communication files to be associated; and a third generation unit. The system is configured to: generate corresponding communication file feature information for each of the aforementioned communication files to be associated, and determine the communication files to be associated and the communication file feature information as display file information; a determining unit is configured to determine the obtained display file information as a display file information set; a sending unit is configured to send the recommended association type information, the recommended association communication file set, and the display file information set to the client for the client to display the recommended association type information, the recommended association communication file set, and the display file information set; a fourth generating unit is configured to, in response to receiving the filtering association interaction information sent by the client, generate communication file association parsing information based on the filtering association interaction information and the aforementioned communication files to be associated; and a second sending unit is configured to send the communication file association parsing information to the client.
[0010] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0011] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0012] The above-described embodiments of this disclosure have the following beneficial effects: the communication file association resolution information generation method of some embodiments of this disclosure improves the user experience. Specifically, the reason for the poor user experience is that users view various communication files, select the communication files to be associated, and then perform cross-file content association and extraction on the selected multiple communication files based on general rules or standards (such as file type, creation time, etc.) to obtain communication file association resolution information. This relies entirely on the user's own selection and association of files, without personalized recommendations based on the user's historical behavior or preferences, which may lead to users spending more time and effort to filter and associate files, resulting in a poor user experience. Based on this, the communication file association resolution information generation method of some embodiments of this disclosure firstly, in response to receiving communication file association resolution task information sent by the client, collects each communication file to be associated corresponding to the above communication file association resolution task information, wherein the above communication file association resolution task information includes a user identifier, and each of the above communication files to be associated corresponds to a service object communicator identifier and a file identifier. Thus, each communication file to be associated can be obtained for generating each recommended associated communication file set. Next, communication file association behavior profile information corresponding to the above user identifier is obtained. This allows us to obtain a profile of communication document association behavior that represents a user's past preferences when associating communication documents. Then, based on this profile, recommended association type information is generated. This allows us to generate personalized recommendations that better suit the user's actual needs, building upon the profile of communication document association behavior representing the user's historical behavior or preferences. Next, based on the recommended association type information, the communication document association behavior profile, and the various communication documents to be associated, we generate sets of recommended communication documents corresponding to the recommended association type information. This generates at least one set of documents that provides users with directly referable association options, forming various sets of recommended communication documents. Then, for each of the communication documents to be associated, we generate corresponding communication document feature information and determine the communication document to be associated and its feature information as display document information. This generates communication document feature information for each communication document to be associated, allowing us to display the main content of the document to the user. Finally, we determine the resulting display document information set. Subsequently, the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information are sent to the aforementioned client so that the client can display the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information.Therefore, the system can display recommended association type information, the aforementioned recommended association communication file sets, and the aforementioned displayed file information set, allowing users to select communication files to be associated and extracted for content association. This generates filtering association interaction information for the selected communication files. Next, in response to receiving the filtering association interaction information from the client, communication file association resolution information is generated based on the filtering association interaction information and the aforementioned communication files. This generates communication file association resolution information corresponding to the selected communication files. Finally, the communication file association resolution information is sent to the client. Because, before the user selects a communication file to be associated and extracted, personalized recommended association type information that better suits the user's actual needs is generated based on the communication file association behavior profile information representing the user's historical behavior or preferences, and at least one set of files providing users with directly referable association options is generated, before the user selects a communication file to be associated and extracted for content association, each recommended association communication file set is formed. The recommended association type information and recommended association communication file sets provide users with a wealth of reference information, helping them to have a more comprehensive understanding of the file content and association possibilities. This assists users in filtering or selecting communication files to be associated and extracted, thus improving the user experience. Attached Figure Description
[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0014] Figure 1 This is a flowchart of some embodiments of the communication document association parsing information generation method according to this disclosure;
[0015] Figure 2 This is a schematic diagram of the structure of some embodiments of the communication document association parsing information generation apparatus according to the present disclosure;
[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] Figure 1 A flow 100 of some embodiments of a communication document association resolution information generation method according to the present disclosure is shown. This communication document association resolution information generation method includes the following steps:
[0024] Step 101: In response to receiving the communication file association and parsing task information sent by the client, collect each communication file to be associated corresponding to the communication file association and parsing task information.
[0025] In some embodiments, the executing entity (e.g., a computing device) of the communication document association and parsing information generation method can, in response to receiving communication document association and parsing task information sent by a client, collect each communication document to be associated corresponding to the aforementioned communication document association and parsing task information. The aforementioned communication document association and parsing task information includes a user identifier, and each of the communication documents to be associated corresponds to a service recipient communicator identifier and a file identifier. The aforementioned communication document association and parsing task information can be information representing the execution of the communication document association and parsing task. The aforementioned association and parsing task can associate and parse each communication document corresponding to the user identifier to generate a summary of the communication content for each communication document. The aforementioned user identifier can be the identifier of a user participating in communication (e.g., online communication such as conference communication or voice communication) or initiating online communication (e.g., a headhunter user identifier). The aforementioned service recipient communicator identifier can be the identifier of the object with which the user corresponding to the aforementioned user identifier conducts online communication (e.g., an applicant identifier or a recruiter identifier). In practice, the executing entity can collect each communication document corresponding to the user identifier included in the aforementioned communication document association and parsing task information from a preset database as each communication document to be associated. Each of the aforementioned communication files to be associated can be an electronic file used to transmit information and record discussion content (e.g., meeting video, audio, chat log text files). The aforementioned communication files to be associated can be, but are not limited to, one of the following: meeting video files, audio files, or chat log text files. The aforementioned preset database can be a structured data system (e.g., a MySQL database) used to store various communication records (such as meeting videos, audio, chat logs, emails, etc.).
[0026] Step 102: Obtain the communication document-related behavior profile information corresponding to the user identifier.
[0027] In some embodiments, the aforementioned executing entity may obtain communication document-related behavioral profile information associated with the aforementioned user identifier.
[0028] In some optional implementations of certain embodiments, the aforementioned executing entity may obtain the communication document-related behavioral profile information corresponding to the aforementioned user identifier through the following steps:
[0029] The first step is to retrieve the historical association type information set corresponding to the aforementioned user identifier from the preset communication file association log. This preset communication file association log records the association type information when a user associates various communication files. Each historical association type information in the aforementioned historical association type information set can represent the association type of the communication file (e.g., voiceprint homology association type, file content similarity association type, service object same identifier association type). The voiceprint homology association type indicates that communication files representing communicators with similar voice characteristics are associated. The file content similarity association type indicates that communication files with similar communication content are associated. The service object same identifier association type indicates that communication files with the same service object identifier are associated.
[0030] The second step is to determine the target time point as the time point when the communication file sent by the client is associated with the parsing task information.
[0031] The third step involves obtaining user click and query records for each communication file to be associated within a preset time period prior to the target time point. Each click record includes the click time, and each click record corresponds to one of the communication files to be associated. The click records represent user clicks on the communication files to be associated. The query records represent user queries for the communication files to be associated within the preset time period prior to the target time point. The query records include various query records. Each query record includes a file identifier or a service recipient / communicator identifier. The file identifier can be the filename of the communication file to be associated.
[0032] The fourth step is to identify the above-mentioned historical association type information set, the above-mentioned click record information and query information as communication document association behavior profile information.
[0033] Step 103: Based on the communication document association behavior profile information, generate recommended association type information.
[0034] In some embodiments, the aforementioned executing entity may generate recommended association type information based on the aforementioned communication document-related behavioral profile information.
[0035] In some optional implementations of certain embodiments, the aforementioned executing entity can generate recommended association type information based on the aforementioned communication document-related behavioral profile information through the following steps:
[0036] The first step is to cluster the historical association type information set included in the aforementioned communication document association behavior profile information to obtain various historical association type information groups. Each of these historical association type information groups contains identical historical association type information. In practice, the executing entity can group identical historical association type information from the historical association type information set to obtain various historical association type information groups. For example, the aforementioned historical association type information set could be “{Voiceprint same-origin association type information, Service object same-identity association type information, Document content similarity association type information, Voiceprint same-origin association type information, Voiceprint same-origin association type information, Document content similarity association type information}”. Each historical association type information group could be “{Voiceprint same-origin association type information, Voiceprint same-origin association type information, Voiceprint same-origin association type information}, {Service object same-identity association type information}, {Document content similarity association type information, Document content similarity association type information}”.
[0037] The second step is to identify the historical association type information group with the largest number of historical association type information groups as the target historical association type information group. For example, the target historical association type information group can be "{voiceprint homology association type information, voiceprint homology association type information, voiceprint homology association type information}".
[0038] The third step is to identify one of the historical association type information from the aforementioned target historical association type information group as the recommended association type information.
[0039] Step 104: Based on the recommendation association type information, the communication file association behavior profile information, and each communication file to be associated, generate each set of recommendation association communication files corresponding to the recommendation association type information.
[0040] In some embodiments, the executing entity may generate a set of recommended associated communication files corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and the various communication files to be associated.
[0041] In some optional implementations of certain embodiments, the execution entity may generate a set of recommendation-related communication files corresponding to the recommendation-related type information by means of the following steps: based on the recommendation-related type information, the communication file-related behavior profile information, and the various communication files to be associated.
[0042] The first step is to identify the click records included in the aforementioned communication document-related behavioral profile information as individual target click records. Each target click record corresponds to one of the aforementioned communication documents to be associated.
[0043] The second step is to identify the query information included in the aforementioned communication document-related behavioral profile information as the target query information. This target query information includes at least one service recipient communicator identifier.
[0044] The third step is to identify the aforementioned communication documents to be associated as a candidate set of communication documents to be associated.
[0045] The fourth step is to determine the candidate communication files to be associated as the candidate communication file set, which corresponds to the click record information of each target and the communicator identifier of at least one service object.
[0046] Fifth, in response to determining that the above recommended association type information is voiceprint homology association type information, the following steps are performed:
[0047] The first sub-step involves identifying each candidate communication file of type audio file in the aforementioned candidate communication file set as a separate audio communication file to be associated.
[0048] The second sub-step involves performing audio clustering processing on the aforementioned audio communication files to be associated, obtaining at least one set of audio communication files to be associated as each recommended set of associated communication files. In practice, firstly, the executing entity can use the x-vector model to extract audio features from each audio communication file to be associated, obtaining various audio feature information. Each audio feature information can represent a feature vector of the voice features of each speaker in one of the audio communication files to be associated. Then, the executing entity can use a hierarchical clustering algorithm to cluster the various audio feature information, obtaining at least one group of audio feature information. Next, for each of the at least one group of audio feature information, the executing entity can determine at least one audio communication file corresponding to the aforementioned audio feature information group as the set of audio communication files to be associated.
[0049] Step 6: In response to determining that the above-mentioned recommended association type information is file content similarity association type information, based on the above-mentioned communication files to be associated and the above-mentioned candidate communication file sets to be associated, each recommended association communication file set is generated. The above-mentioned file content similarity association type information can represent an association type that associates at least one communication file according to content similarity. The above-mentioned voiceprint homology association type information can represent an association type that associates at least one communication file according to similar voice features. For example, communication file 1 contains the voices of person A and user B. Communication file 2 contains the voices of person A and user B, then communication file 1 and communication file 2 can be associated to obtain a summary of the communication content of communication files 1 and 2.
[0050] Step 7: In response to determining that the above-mentioned recommended association type information is service object same-identity association type information, based on the above-mentioned communication files to be associated and the above-mentioned candidate communication file sets to be associated, generate each recommended association communication file set. The above-mentioned service object same-identity association type information can indicate that communication files with the same service object identifier will be associated.
[0051] In some optional implementations of certain embodiments, the execution entity may generate various recommended association communication file sets based on the various communication files to be associated and the candidate communication file sets to be associated in response to determining that the recommended association type information is file content similarity association type information:
[0052] First, in response to determining that the above recommended association type information is file content similarity association type information, for each candidate communication file in the above candidate communication file set to be associated, the following steps are performed:
[0053] The first sub-step involves identifying at least one target communication file among the aforementioned communication files to be associated, whose communication content similarity to the candidate communication files to be associated is greater than a preset content similarity. In practice, firstly, the executing entity can input the candidate communication files to be associated (e.g., text, audio, image / video files) into the Perceiver large model to obtain candidate file content information corresponding to the aforementioned candidate communication files to be associated. This candidate file content information can represent the file content of the candidate communication files to be associated, and can be represented by vectors or text. Then, for each communication file to be associated, the executing entity can input the communication file to be associated into the Perceiver large model to obtain the corresponding file content information as the communication file content information to be associated. This communication file content information can represent the file content of the communication file to be associated, and can be represented by vectors or text. Finally, the executing entity can determine the similarity between each communication file content information and the candidate file content information as the communication content similarity. The aforementioned similarity can be represented by the cosine similarity between the content information of the communication file to be associated and the content information of the candidate files. Each communication content similarity corresponds to one of the communication files to be associated. Finally, the executing entity can identify at least one communication file among the communication files to be associated whose communication content similarity with the aforementioned candidate communication files is greater than a preset content similarity as at least one target communication file to be associated.
[0054] The second sub-step involves identifying the aforementioned candidate communication documents to be associated and at least one target communication document to be associated as the respective recommended communication documents.
[0055] The third sub-step involves defining the identified recommended related communication documents as a set of recommended related communication documents.
[0056] In some optional implementations of certain embodiments, the execution entity may generate various recommended association communication file sets based on the various communication files to be associated and the candidate communication file sets to be associated in response to determining that the recommended association type information is service object and identifier association type information:
[0057] The first step is to determine the target service object communicator identifier for each candidate communication file in the above candidate communication file set.
[0058] The second step is to deduplicate the communicator identifiers of each target service object to obtain deduplicated communicator identifiers for each service object.
[0059] The third step is to identify each deduplication service object communicator identifier in the above deduplication service object communicator identifiers, and determine at least one communication file in each of the above communication files to be associated that corresponds to the above deduplication service object communicator identifier as the recommended association communication file set.
[0060] Step 105: For each communication file to be associated in each communication file to be associated, generate the corresponding communication file feature information, and determine the communication file to be associated and the communication file feature information to be associated as the display file information.
[0061] In some embodiments, the execution entity can generate corresponding communication file feature information for each of the communication files to be associated, and determine the communication file and its feature information as display file information. In practice, for each communication file to be associated, in response to determining that the communication file is a non-text file, the execution entity can input the communication file into the Perceiver large model to obtain the file content represented in text. Then, the execution entity can use keyword extraction technology (e.g., TF-IDF keyword extraction technology) to extract keywords from the text-represented file content as the corresponding communication file feature information. In response to determining that the communication file is a text file, the execution entity can use keyword extraction technology to extract keywords from the communication file as the corresponding communication file feature information.
[0062] Step 106: Determine the obtained display file information as a display file information set.
[0063] In some embodiments, the aforementioned executing entity may determine the obtained display file information as a display file information set.
[0064] Step 107: Send the recommended association type information, each recommended association communication file set, and the display file information set to the client so that the client can display the recommended association type information, each recommended association communication file set, and the display file information set.
[0065] In some embodiments, the executing entity may send the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information to the client, so that the client can display the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information. The client may identify the terminal device (e.g., mobile phone, computer, etc.) used by the user.
[0066] Step 108: In response to receiving the filtering association interaction information sent by the client, generate communication file association parsing information based on the filtering association interaction information and each communication file to be associated.
[0067] In some embodiments, the executing entity may, in response to receiving the filtering association interaction information sent by the client, generate communication file association resolution information based on the filtering association interaction information and the various communication files to be associated. The filtering association interaction information includes association type information and file identifiers. The association type information may be information representing the association type selected by the user through interface interaction (e.g., voiceprint homology association type information, file content similarity association type information, service object same identifier association type information). The file identifiers included in the association type information in the filtering association interaction information may be the file identifiers corresponding to the various communication files to be associated selected in the displayed file information selected by the user through interface interaction.
[0068] In the process of adopting technical solutions to address the problems mentioned in the background section, the following issues often arise:
[0069] When associating audio communication files, the common practice is to convert each file into a text file, extract key content from each text file, and then merge or splice the extracted key content to obtain the communication file association resolution information. However, this method of converting audio to text loses the voiceprint feature information in the audio, making it impossible to distinguish between different speakers. In multi-person communication scenarios, the viewpoints and content of different speakers may be intertwined. Relying solely on text content makes it difficult to accurately understand each speaker's intentions and viewpoints, leading to reduced accuracy in association resolution and a chaotic, disorganized association resolution information, resulting in a poor user experience.
[0070] In some optional implementations of certain embodiments, the aforementioned execution entity can generate communication file association resolution information based on the aforementioned filtered associated interaction information and the aforementioned communication files to be associated through the following steps:
[0071] The first step is to identify the file identifier in each of the file identifiers included in the above-mentioned filtering and association interaction information as the associated communication file.
[0072] The second step, in response to the determination that all the identified associated communication files are audio files, is to perform the following voiceprint homology association processing on each associated communication file:
[0073] The identified associated communication files are processed through audio segmentation to obtain initial audio segments. In practice, the aforementioned execution entity can use the VAD speech segmentation algorithm to perform audio segmentation on the associated communication files to obtain initial audio segments. Here, the file type of the aforementioned associated communication files is audio files.
[0074] The first sub-step involves removing background noise from each of the initial audio segments to obtain individual audio segments. In practice, for each of the initial audio segments, the execution entity can use spectral subtraction to remove background noise, resulting in an audio segment with background noise removed.
[0075] The second sub-step involves extracting voiceprint features from each of the aforementioned audio segments to obtain corresponding voiceprint feature information for each audio segment. Each audio segment corresponds to one of the voiceprint feature information entries. In practice, for each audio segment, the executing entity can input the audio segment into a pre-trained x-vector model to obtain voiceprint feature information. This voiceprint feature information can be voiceprint features extracted from the audio segments, and these voiceprint features can be represented by feature vectors.
[0076] The third sub-step involves clustering the aforementioned voiceprint feature information to obtain various clustered voiceprint feature information groups. In practice, the executing entity can employ a hierarchical clustering algorithm to cluster the various voiceprint feature information information to obtain various clustered voiceprint feature information groups.
[0077] The fourth sub-step involves generating groups of audio segments with the same source voiceprint feature information based on the aforementioned clustered voiceprint feature information groups. In practice, for each clustered voiceprint feature information group, the executing entity can identify the audio segments in each audio segment that correspond to the voiceprint feature information in the aforementioned clustered voiceprint feature information group as groups of audio segments with the same source voiceprint feature information.
[0078] The fifth sub-step involves performing transcription and summary extraction on each of the aforementioned homologous voiceprint audio segment groups to obtain communication document association parsing sub-information. In practice, for each homologous voiceprint audio segment in the aforementioned homologous voiceprint audio segment groups, the executing entity can use Automatic Speech Recognition (ASR) technology to transcribe the language content of the homologous voiceprint audio segment into text information. Then, the executing entity can input the transcribed text information into a pre-trained generative summarization model (e.g., the PALM Chinese summarization generation model) to obtain communication document association parsing sub-information. This communication document association parsing sub-information can be concise and summarized text information obtained by summarizing the transcribed text information.
[0079] The sixth sub-step involves determining the obtained file association resolution sub-information as the communication file association resolution information.
[0080] The first to the sixth sub-step of the second step of the above technical solution and its related content serve as an inventive point of this disclosure, solving the technical problem of "poor user experience". Factors leading to a poor user experience often include: when associating audio communication files, the common practice is to convert each communication file to text content, extract key content from each text file, and then fuse or splice the extracted key content to obtain the communication file association parsing information. However, this method of converting audio to text loses the voiceprint feature information in the audio, making it impossible to distinguish the speech of different speakers. In multi-person communication scenarios, the viewpoints and speech content of different speakers may be intertwined, making it difficult to accurately understand the intentions and viewpoints of each speaker based solely on text content. This leads to reduced accuracy in association parsing and a chaotic content of the parsed association information, resulting in a poor user experience. Solving these factors can improve the user experience. To achieve this, in the first step, for each file identifier included in the above-mentioned filtering of association interaction information, the communication file corresponding to the aforementioned file identifier among the communication files to be associated is identified as the associated communication file. Therefore, the associated communication files selected by the user for association can be determined. The second step, in response to the determination that all the identified associated communication files are audio files, performs the following voiceprint homology association processing on each associated communication file: First, audio segmentation is performed on each identified associated communication file to obtain initial audio segments. Thus, when the association type information is voiceprint homology association type information, audio segmentation is performed on each associated communication file selected by the user and of audio file type to obtain initial audio segments. Next, background noise removal processing is performed on each of the initial audio segments to obtain individual audio segments. This removes background noise from each initial audio segment. Then, voiceprint feature extraction processing is performed on each of the audio segments to obtain voiceprint feature information corresponding to each audio segment, wherein each audio segment corresponds to one voiceprint feature information in each of the voiceprint feature information. This yields voiceprint feature information used to generate each clustered voiceprint feature information group. Then, the voiceprint feature information is clustered to obtain each clustered voiceprint feature information group. Therefore, we can obtain various clustered voiceprint feature information groups used to generate different groups of audio segments with the same source voiceprint. Then, based on these clustered voiceprint feature information groups, we generate different groups of audio segments with the same source voiceprint, thus distinguishing the speech of different speakers. Consequently, audio segments with similar voiceprint features can be grouped together to form groups of audio segments with the same source voiceprint (i.e., audio segments belonging to the same speaker are grouped together to form groups of audio segments with the same source voiceprint), resulting in different groups of audio segments with the same source voiceprint.Subsequently, for each of the aforementioned homologous voiceprint audio segment groups, transcription and abstract extraction are performed to obtain communication document association parsing sub-information. Thus, based on the transcription and abstract extraction of homologous voiceprint audio segment groups, the communication document association parsing sub-information contains the viewpoints and statements of the same speaker, reducing the possibility of overlapping viewpoints and statements from different speakers. The obtained document association parsing sub-information is then identified as the communication document association parsing information. Therefore, communication document association parsing information can be generated through homologous voiceprint association processing. This information includes document association parsing sub-information that distinguishes the communication content of different speakers, reducing the possibility of overlapping viewpoints and statements from different speakers in multi-person communication scenarios. It also reduces the chaotic content of the association parsing information caused by directly merging or splicing extracted key content, improving the user experience.
[0081] The third step involves, in response to the determination that the aforementioned association type information is file content similarity association type information, inputting the aforementioned preset file content similarity association instruction information and the determined associated communication files into a pre-trained communication file association resolution information generation model to obtain communication file association resolution information. In practice, firstly, the executing entity can call a preset speech recognition toolkit (e.g., the Kaldi speech recognition toolkit) to convert the determined associated communication files (each with an audio file type) into text files to update the associated communication files, obtaining the files in the updated associated communication files. Then, the executing entity can input the preset file content similarity association instruction information and the updated associated communication files into the communication file association resolution information generation model to obtain communication file association resolution information. The aforementioned preset file content similarity association instruction information can be a preset instruction or prompt word that performs cross-file content similarity analysis on the files in the updated associated communication files based on their file content. For example, the preset file content similarity association instruction information could be "Perform similarity analysis on the input updated associated communication files and generate a summary report accordingly." The aforementioned communication document association resolution information can refer to the association information (e.g., common themes, similar viewpoints, or related trends) extracted from various updated related communication documents through content similarity analysis. The generation model for this communication document association resolution information can be a large language model (e.g., Doubao, Deepseek).
[0082] The fourth step involves determining, in response to the identification of the aforementioned association type information as service object and identifier association type information, identifying the service object communicator identifiers corresponding to each associated communication file, and generating communication file association resolution information based on the service object communicator identifiers and the associated communication files. Each service object communicator identifier corresponds to one associated communication file among the associated communication files. In practice, firstly, the executing entity can determine each service object communicator identifier as a service object communicator identifier set. Then, it can determine each associated communication file as an associated communication file set. Next, the executing entity can perform deduplication on the service object communicator identifier set to obtain a deduplicated service object communicator identifier set. Then, for each deduplicated service object communicator identifier in the deduplicated service object communicator identifier set, the executing entity can determine each associated communication file in the associated communication file set corresponding to the deduplicated service object communicator identifier as a target associated communication file. Afterward, the executing entity can input each target associated communication file and preset file association instruction information into the communication file association resolution information generation model to obtain model output information corresponding to the deduplicated service object communicator identifiers. Next, the model output information is identified as the communication document association parsing sub-information. Finally, the aforementioned executing entity can define the identified communication document association parsing sub-information as the communication document association parsing information. The aforementioned preset document association instruction information can be an instruction or prompt that integrates the file content corresponding to each target associated communication document. The aforementioned communication document association parsing sub-information can be a comprehensive analysis report or summary obtained after integrating the file content corresponding to each target associated communication document.
[0083] In some optional implementations of certain embodiments, after sending the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information to the aforementioned client for the client to display the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information, the aforementioned execution entity may further perform the following steps:
[0084] The first step involves generating pre-loaded communication file association resolution information corresponding to each of the aforementioned recommended association type information and recommended association communication file sets. Each recommended association communication file set corresponds to one of the pre-loaded communication file association resolution information sets. In practice, for each recommended association communication file set, the following pre-association processing is performed: First sub-step: In response to determining that the recommended association type information is a voiceprint homology association type, voiceprint homology association processing is performed on each recommended association communication file in the recommended association communication file set to obtain pre-loaded communication file association resolution information. Second sub-step: In response to determining that the recommended association type information is a file content similarity association type, preset file content similarity association instruction information and each recommended association communication file in the recommended association communication file set are input into a pre-trained communication file association resolution information generation model to obtain pre-loaded communication file association resolution information. The third sub-step, in response to determining that the above-mentioned recommendation association type information is service object same identifier association type information, determines the service object communicator identifier corresponding to each recommendation association communication file in each recommendation association communication file set, and generates preloaded communication file association resolution information based on the above-mentioned service object communicator identifier and the above-mentioned recommendation association communication file.
[0085] The second step is to determine the association resolution information of each of the above pre-loaded communication files as the association resolution information set of the communication files to be transmitted.
[0086] The third step is to obtain the pre-uploaded audio data corresponding to the user identifier. In practice, the aforementioned execution entity can obtain the pre-uploaded audio data corresponding to the user identifier from a preset file. This pre-uploaded audio data can represent pre-collected speech segments of the user.
[0087] The fourth step is to determine the upload time point corresponding to the above-mentioned pre-uploaded audio data as the target upload time point.
[0088] Fifth step: In response to determining that the target upload time point is within a preset daytime period, background noise removal processing is performed on the pre-uploaded audio data to obtain pre-processed upload audio data.
[0089] The sixth step involves pre-emphasizing the pre-processed uploaded audio data to obtain enhanced audio data. In practice, the executing entity can use high-pass filtering technology to pre-emphasize the pre-processed uploaded audio data to obtain enhanced audio data.
[0090] Step 7: Perform frame-segmentation and windowing processing on the enhanced audio data to obtain an audio frame data sequence. In practice, the execution entity can use frame-segmentation and windowing technology to process the enhanced audio data and obtain an audio frame data sequence. Each audio frame in the audio frame data sequence can be the data corresponding to an audio segment in the enhanced audio data. The audio frame data sequence can represent the audio signal segment sequence obtained by dividing the continuous enhanced audio signal into short, fixed-length discrete segments (i.e., frames) using frame-segmentation and windowing technology, and applying a window function to each frame signal.
[0091] Step 8: Perform sound feature extraction processing on the aforementioned audio frame data sequence to obtain sound feature vector information corresponding to the aforementioned user identifier. In practice, for each audio frame data in the aforementioned audio frame data sequence, the aforementioned execution entity can use Mel-frequency cepstral coefficient (MFCC) technology to perform sound feature extraction processing on the aforementioned audio frame data, obtaining the MFCC feature vector corresponding to the aforementioned audio frame data as the sound feature vector sub-information. Then, the obtained sound feature vector sub-information is averaged dimension by dimension to obtain the MFCC feature vector, and the average value of each dimension is taken as the sound feature vector information.
[0092] Step nine: Based on the aforementioned sound feature vector information, generate encryption key information. In practice, the executing entity can input the sound feature vector information into a preset fuzzy extractor to obtain the encryption key. The obtained encryption key is then designated as the encryption key information.
[0093] Step 10: Based on the aforementioned encryption key information, encrypt the aforementioned communication file association resolution information set to be transmitted to obtain an encrypted communication file association resolution information set. In practice, the executing entity can use a preset encryption function (e.g., AES encryption function) to encrypt the communication file association resolution information set to be transmitted according to the encryption key information, thereby obtaining the encrypted communication file association resolution information set.
[0094] Step 11: Send the encrypted transmission file association and parsing information set to the client, so that the client can be configured to perform the following offline processing steps:
[0095] The first sub-step is to cache the encrypted transmission file association parsing information in memory.
[0096] The second sub-step, in response to detecting a user's association resolution selection operation with one of the displayed recommended associated communication sets, involves collecting the user's voice to obtain verification audio data. In practice, the client can use a microphone to collect the user's voice to obtain the verification audio data. This verification audio data can be the collected audio data of the user's voice.
[0097] The third sub-step involves extracting sound features from the aforementioned verification audio data to obtain verification sound feature vector information. In practice, the client can use Mel-frequency cepstral coefficients (MFCCs) to extract sound features from the verification audio data, obtaining the MFCC feature vector corresponding to the verification audio data as the verification sound feature vector information.
[0098] The fourth sub-step involves generating decryption key information based on the aforementioned verified sound feature vector information. In practice, the client can input the verified sound feature vector information into the aforementioned preset fuzzy extractor to obtain the decryption key as the decryption key information.
[0099] The fifth sub-step involves decrypting the encrypted communication file association resolution information set in memory based on the aforementioned decryption key information, and displaying the communication file association resolution information set corresponding to the recommended associated communication file set after decryption. In practice, the client can use a preset decryption algorithm to decrypt the encrypted communication file association resolution information set in memory based on the decryption key information to obtain the decrypted communication file association resolution information set.
[0100] The above-described technical solution and related content, as an inventive point of this disclosure, solve the technical problem of "long waiting time for users to obtain communication file association resolution information corresponding to recommended associated communication file sets and low security of communication file association resolution information transmission." The factors leading to long waiting times for users to obtain communication file association resolution information corresponding to recommended associated communication file sets and low security of communication file association resolution information transmission are often as follows: When a user clicks on a recommended associated communication file set in the client, they usually want to quickly obtain the communication file association resolution information corresponding to the recommended associated communication file set. If the communication file association resolution task information is sent in real time, the waiting time may be too long due to network latency, offline status, server load, etc., affecting the user experience. At the same time, the pre-generated pre-loaded communication file association resolution information for each recommended associated communication file set usually contains a large amount of user privacy data. If the pre-generated pre-loaded communication file association resolution information is directly sent to the client for caching, privacy data is easily leaked during the transmission process, resulting in low security of communication file association resolution information transmission. Solving the above factors can improve the user experience and enhance the security of communication file association resolution information transmission. To achieve this effect, firstly, based on the aforementioned recommended association type information and the aforementioned recommended association communication file sets, preloaded communication file association resolution information corresponding to each of the aforementioned recommended association communication file sets is generated. Each recommended association communication file set corresponds to one preloaded communication file association resolution information in the aforementioned preloaded communication file association resolution information. Therefore, the preloaded communication file association resolution information corresponding to each of the aforementioned recommended association communication file sets can be generated in advance before the user might click on the recommended file set. Then, the aforementioned preloaded communication file association resolution information is determined as the set of communication file association resolution information to be transmitted. Next, pre-uploaded audio data corresponding to the user identifier is obtained. This allows the acquisition of pre-uploaded audio data used to generate encryption key information. Then, the upload time point corresponding to the aforementioned pre-uploaded audio data is determined as the target upload time point. Then, in response to determining that the target upload time point is within a preset daytime period, background noise removal processing is performed on the aforementioned pre-uploaded audio data to obtain preprocessed uploaded audio data. Therefore, when the pre-uploaded audio data is collected during the day, background noise removal processing can be performed on the pre-uploaded audio data to reduce background noise. Then, the preprocessed uploaded audio data is pre-emphasized to obtain enhanced audio data. Next, the enhanced audio data is subjected to frame-by-frame windowing to obtain an audio frame data sequence. Then, sound feature extraction processing is performed on the audio frame data sequence to obtain sound feature vector information corresponding to the user identifier. Finally, based on the sound feature vector information, encryption key information is generated.Therefore, through the above steps, background noise removal, pre-emphasis, and voice feature extraction can be performed on the pre-uploaded audio data to generate the user's voice feature vector information. Based on the user's voice feature vector information, encryption key information is generated for encrypting the communication file association parsing information set to be transmitted. Then, based on the encryption key information, the communication file association parsing information set to be transmitted is encrypted to obtain an encrypted communication file association parsing information set. This encrypts the communication file association parsing information set to be transmitted, ensuring that privacy data is not easily leaked when the encrypted communication file association parsing information set is subsequently sent to the client, thus improving the security of the communication file association parsing information transmission. The encrypted communication file association parsing information set is sent to the client, allowing the client to be configured to perform the following offline processing steps: First sub-step: The encrypted communication file association parsing information is cached in memory. Second sub-step: In response to detecting a user's association parsing selection operation with one of the displayed recommended communication file sets, the user's voice is collected to obtain verification audio data. The third sub-step involves extracting sound features from the aforementioned verification audio data to obtain verification sound feature vector information. This provides the verification sound feature vector information used to generate the decryption key. The fifth sub-step, based on the decryption key information, decrypts the encrypted communication file association resolution information set in memory and displays the pre-loaded communication file association resolution information corresponding to the recommended communication file set within the decrypted communication file association resolution information set. Thus, the encrypted communication file association resolution information set can be stored in the client's memory through the aforementioned offline processing steps. This utilizes the high-speed access characteristics of memory (nanosecond level) to replace network requests, reducing the waiting time for users to wait for the pre-loaded communication file association resolution information corresponding to the recommended communication file set due to network latency, offline status, server load, or other reasons.
[0101] Step 109: Send the communication document association and parsing information to the client.
[0102] In some embodiments, the aforementioned executing entity may send the aforementioned communication document association parsing information to the aforementioned client.
[0103] The above-described embodiments of this disclosure have the following beneficial effects: the communication file association resolution information generation method of some embodiments of this disclosure improves the user experience. Specifically, the reason for the poor user experience is that users view various communication files, select the communication files to be associated, and then perform cross-file content association and extraction on the selected multiple communication files based on general rules or standards (such as file type, creation time, etc.) to obtain communication file association resolution information. This relies entirely on the user's own selection and association of files, without personalized recommendations based on the user's historical behavior or preferences, which may lead to users spending more time and effort to filter and associate files, resulting in a poor user experience. Based on this, the communication file association resolution information generation method of some embodiments of this disclosure firstly, in response to receiving communication file association resolution task information sent by the client, collects each communication file to be associated corresponding to the above communication file association resolution task information, wherein the above communication file association resolution task information includes a user identifier, and each of the above communication files to be associated corresponds to a service object communicator identifier and a file identifier. Thus, each communication file to be associated can be obtained for generating each recommended associated communication file set. Next, communication file association behavior profile information corresponding to the above user identifier is obtained. This allows us to obtain a profile of communication document association behavior that represents a user's past preferences when associating communication documents. Then, based on this profile, recommended association type information is generated. This allows us to generate personalized recommendations that better suit the user's actual needs, building upon the profile of communication document association behavior representing the user's historical behavior or preferences. Next, based on the recommended association type information, the communication document association behavior profile, and the various communication documents to be associated, we generate sets of recommended communication documents corresponding to the recommended association type information. This generates at least one set of documents that provides users with directly referable association options, forming various sets of recommended communication documents. Then, for each of the communication documents to be associated, we generate corresponding communication document feature information and determine the communication document to be associated and its feature information as display document information. This generates communication document feature information for each communication document to be associated, allowing us to display the main content of the document to the user. Finally, we determine the resulting display document information set. Subsequently, the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information are sent to the aforementioned client so that the client can display the aforementioned recommendation association type information, the aforementioned sets of recommendation association communication documents, and the aforementioned set of display document information.Therefore, the system can display recommended association type information, the aforementioned recommended association communication file sets, and the aforementioned displayed file information set, allowing users to select communication files to be associated and extracted for content association. This generates filtering association interaction information for the selected communication files. Next, in response to receiving the filtering association interaction information from the client, communication file association resolution information is generated based on the filtering association interaction information and the aforementioned communication files. This generates communication file association resolution information corresponding to the selected communication files. Finally, the communication file association resolution information is sent to the client. Because, before the user selects a communication file to be associated and extracted, personalized recommended association type information that better suits the user's actual needs is generated based on the communication file association behavior profile information representing the user's historical behavior or preferences, and at least one set of files providing users with directly referable association options is generated, before the user selects a communication file to be associated and extracted for content association, each recommended association communication file set is formed. The recommended association type information and recommended association communication file sets provide users with a wealth of reference information, helping them to have a more comprehensive understanding of the file content and association possibilities. This assists users in filtering or selecting communication files to be associated and extracted, thus improving the user experience.
[0104] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a communication document association parsing information generation apparatus, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0105] like Figure 2As shown, the communication document association parsing information generation device 200 in some embodiments includes: a collection unit 201, an acquisition unit 202, a first generation unit 203, a second generation unit 204, a third generation unit 205, a determination unit 206, a sending unit 207, a fourth generation unit 208, and a second sending unit 209. The acquisition unit 201 is configured to, in response to receiving communication file association parsing task information sent by the client, acquire each communication file to be associated corresponding to the aforementioned communication file association parsing task information. The aforementioned communication file association parsing task information includes a user identifier, and each communication file to be associated corresponds to a service object communicator identifier and a file identifier. The acquisition unit 202 is configured to acquire communication file association behavior profile information corresponding to the aforementioned user identifier. The first generation unit 203 is configured to generate recommended association type information based on the aforementioned communication file association behavior profile information. The second generation unit 204 is configured to generate a set of recommended associated communication files corresponding to the recommended association type information based on the aforementioned recommended association type information, the aforementioned communication file association behavior profile information, and the aforementioned communication files to be associated. The third generation unit 205 is configured to, for each of the aforementioned communication files to be associated... For each communication file to be associated, generate corresponding communication file feature information, and determine the communication file to be associated and the communication file feature information as display file information; the determining unit 206 is configured to determine the obtained display file information as a display file information set; the sending unit 207 is configured to send the recommended association type information, the recommended association communication file sets, and the display file information set to the client so that the client can display the recommended association type information, the recommended association communication file sets, and the display file information set; the fourth generating unit 208 is configured to generate communication file association parsing information based on the filtered association interaction information and the communication files to be associated in response to receiving the filtered association interaction information sent by the client; the second sending unit 209 is configured to send the communication file association parsing information to the client.
[0106] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0107] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0108] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0109] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0110] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0111] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0112] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0113] The computer-readable medium may be contained within an electronic device or may exist independently, not assembled into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: respond to receiving communication file association resolution task information sent by a client; collect each communication file to be associated corresponding to the communication file association resolution task information, wherein the communication file association resolution task information includes a user identifier, and each of the communication files to be associated corresponds to a service recipient communicator identifier and a file identifier; obtain communication file association behavior profile information corresponding to the user identifier; generate recommended association type information based on the communication file association behavior profile information; and generate each recommended associated communication file corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and the communication files to be associated. For each of the aforementioned communication files to be associated, generate corresponding communication file feature information, and determine the communication files to be associated and their feature information as display file information; determine the obtained display file information as a display file information set; send the aforementioned recommended association type information, the aforementioned recommended association communication file sets, and the aforementioned display file information set to the aforementioned client for the client to display the aforementioned recommended association type information, the aforementioned recommended association communication file sets, and the aforementioned display file information set; in response to receiving the filtering association interaction information sent by the aforementioned client, generate communication file association parsing information based on the aforementioned filtering association interaction information and the aforementioned communication files to be associated; send the aforementioned communication file association parsing information to the aforementioned client.
[0114] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0116] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a collection unit, an acquisition unit, a first generation unit, a second generation unit, a third generation unit, a determination unit, a sending unit, a fourth generation unit, and a second sending unit. The names of these units do not necessarily limit the specific unit; for example, the acquisition unit may also be described as "a unit that acquires communication document-related behavioral profile information corresponding to the aforementioned user identifier."
[0117] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0118] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for generating communication document association and parsing information, comprising: In response to receiving communication file association and parsing task information sent by the client, each communication file to be associated corresponding to the communication file association and parsing task information is collected. The communication file association and parsing task information includes a user identifier, and each communication file to be associated corresponds to a service object communicator identifier and a file identifier. Obtain the behavioral profile information associated with the communication documents corresponding to the user identifier; Based on the communication document association behavior profile information, recommended association type information is generated, wherein generating the recommended association type information based on the communication document association behavior profile information includes: Clustering is performed on the historical association type information set included in the communication document association behavior profile information to obtain various historical association type information groups, wherein the historical association type information in each historical association type information group is the same; The historical association type information group with the largest number of historical association type information groups is determined as the target historical association type information group. One of the historical association type information in the target historical association type information group is identified as the recommended association type information; Based on the recommended association type information, the communication file association behavior profile information, and each communication file to be associated, generate a set of recommended association communication files corresponding to the recommended association type information; For each of the communication files to be associated, generate corresponding communication file feature information, and determine the communication file to be associated and the communication file feature information as display file information. The obtained information from each display file is defined as the display file information set; The recommended association type information, the various recommended association communication file sets, and the display file information set are sent to the client so that the client can display the recommended association type information, the various recommended association communication file sets, and the display file information set; In response to receiving the filtering association interaction information sent by the client, communication file association parsing information is generated based on the filtering association interaction information and each communication file to be associated; The communication document association and parsing information is sent to the client.
2. The method according to claim 1, wherein, The step of obtaining the communication document-related behavior profile information corresponding to the user identifier includes: Obtain the set of historical association type information corresponding to the user identifier from the preset communication document association log; The time point at which the communication documents sent by the client are associated with and parsed to determine the task information is the target time point; Obtain user click records and query information for each communication file to be associated within a preset time period before the target time point; The historical association type information set, the various click record information, and the query information are identified as communication document association behavior profile information.
3. The method according to claim 1, wherein, The recommended association type information is one of the following: voiceprint homology association type information, file content similarity association type information, service object same identifier association type information. The step of generating each recommended association communication file set corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and each communication file to be associated includes: Each click record information included in the communication file association behavior profile information is determined as a target click record information, wherein each target click record information corresponds to one of the communication files to be associated. The query information included in the communication document-related behavioral profile information is determined as the target query information, wherein the target query information includes at least one service recipient communicator identifier; Each of the communication files to be associated is identified as a candidate set of communication files to be associated. The candidate communication files to be associated in the candidate communication file set that correspond to the target click record information and the at least one service object communicator identifier are determined as the candidate communication file set to be associated. In response to determining that the recommended association type information is voiceprint homology association type information, the following steps are performed: Each candidate communication file whose file type is audio file is identified as a separate audio communication file to be associated in the candidate communication file set. Perform audio clustering processing on each of the audio communication files to be associated to obtain at least one set of audio communication files to be associated as each set of recommended associated communication files; In response to determining that the recommended association type information is file content similarity association type information, each recommended association communication file set is generated based on each communication file to be associated and the candidate communication file set to be associated; In response to determining that the recommended association type information is service object same identifier association type information, each recommended association communication file set is generated based on each communication file to be associated and the candidate communication file set to be associated.
4. The method according to claim 3, wherein, In response to determining that the recommended association type information is file content similarity association type information, each recommended association communication file set is generated based on the each communication file to be associated and the candidate communication file set to be associated, including: In response to determining that the recommended association type information is file content similarity association type information, for each candidate communication file in the candidate communication file set to be associated, the following steps are performed: At least one communication file to be associated with a communication content similarity greater than a preset content similarity with the candidate communication files to be associated is identified as at least one target communication file to be associated. The candidate communication files to be associated and the at least one target communication file to be associated are determined as the recommended communication files; The identified recommendation-related communication documents are defined as the recommendation-related communication document set.
5. The method according to claim 3, wherein, In response to determining that the recommended association type information is service object same identifier association type information, each recommended association communication file set is generated based on each communication file to be associated and the candidate communication file set to be associated, including: For each candidate communication file to be associated in the candidate communication file set, the service object communicator identifier corresponding to the candidate communication file to be associated is determined as the target service object communicator identifier; The obtained communicator identifiers of each target service object are deduplicated to obtain the deduplicated communicator identifiers of each service object. For each deduplication service object communicator identifier in the various deduplication service object communicator identifiers, at least one communication file in each communication file to be associated that corresponds to the deduplication service object communicator identifier is determined as the recommended associated communication file set.
6. A communication document association parsing information generation device, comprising: The acquisition unit is configured to, in response to receiving communication file association and parsing task information sent by the client, acquire each communication file to be associated corresponding to the communication file association and parsing task information, wherein the communication file association and parsing task information includes a user identifier, and each communication file to be associated corresponds to a service object communicator identifier and a file identifier. The acquisition unit is configured to acquire communication file-related behavioral profile information corresponding to the user identifier; The first generation unit is configured to generate recommended association type information based on the communication document association behavior profile information. The generation of recommended association type information based on the communication document association behavior profile information includes: clustering the historical association type information set included in the communication document association behavior profile information to obtain various historical association type information groups, wherein the historical association type information in each historical association type information group is the same; determining the historical association type information group with the largest number of historical association type information in each historical association type information group as the target historical association type information group; and determining one historical association type information from the target historical association type information group as the recommended association type information. The second generation unit is configured to generate a set of recommended associated communication files corresponding to the recommended association type information based on the recommended association type information, the communication file association behavior profile information, and the various communication files to be associated. The third generation unit is configured to generate, for each of the communication files to be associated, feature information of the communication file to be associated corresponding to the communication file to be associated, and to determine the communication file to be associated and the feature information of the communication file to be associated as display file information. The determining unit is configured to determine the obtained display file information as a display file information set; The sending unit is configured to send the recommendation association type information, the various recommendation association communication file sets, and the display file information set to the client, so that the client can display the recommendation association type information, the various recommendation association communication file sets, and the display file information set; The fourth generation unit is configured to, in response to receiving the filtering association interaction information sent by the client, generate communication file association parsing information based on the filtering association interaction information and each communication file to be associated; The second sending unit is configured to send the communication file association parsing information to the client.
7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Data management method and device, equipment and storage medium
CN114201679A