Communication group processing method and device, and computer equipment

By automating the processing of video and keyword features submitted by target individuals, the system can quickly and accurately determine their eligibility to join a chat group, solving the problem of low processing efficiency in existing technologies and achieving efficient chat group management.

CN121456494APending Publication Date: 2026-02-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411044763.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing methods for handling chat groups, such as relying on users' historical behavior data or manual review, result in low processing efficiency and make it difficult to quickly and accurately identify users who are interested in a particular chat group topic.

Method used

By acquiring the visual and keyword features corresponding to the videos submitted by the target object, determining the quality score of the video based on the visual features, calculating the similarity between the visual features and the preset visual features of the communication group, and the similarity between the keyword features and the preset word features of the communication group, the admission information for the target object to join the communication group is determined comprehensively.

Benefits of technology

It enables the rapid and accurate determination of admission information in the face of a large number of applicants, improves the processing efficiency of communication groups, reduces the subjectivity and inconsistency of manual review, and improves the processing efficiency and objectivity of admission information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456494A_ABST
    Figure CN121456494A_ABST
Patent Text Reader

Abstract

The invention relates to a communication group processing method and device, computer equipment, a storage medium and a computer program product. The method can be applied to an application scene in which a vehicle-mounted terminal, a cloud server or other equipment interacts with an application program. The method comprises the following steps: acquiring visual features and keyword features corresponding to a video submitted by a target object; determining a quality score of the video based on the visual features; determining a visual similarity between the visual feature and a preset visual feature of the communication group; determining a keyword similarity between the keyword feature and a preset word feature of the communication group; and based on the video quality score, the visual similarity and the keyword similarity, determining admission information of the target object joining the communication group. By adopting the method, the processing efficiency of the communication group can be effectively improved, and convenience is brought to users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to an exchange group processing method and device, computer equipment, storage medium and computer program product. BACKGROUND

[0002] With the development of computer technology and Internet technology, the advent of the 5G era, the advent of the Internet has brought great convenience to modern life, and more and more users can share information, emotional communication, business processing, etc. through the use of different multimedia application software, which brings convenience to users. For example, with the continuous popularization of mobile terminals and the acceleration of network access speed, short videos gradually gain the favor of users due to their short, fast and large flow characteristics. In different social media (or short video) platforms at home and abroad, the platform will recommend relevant exchange groups according to the interests and historical behaviors of users to meet the personalized needs of users.

[0003] However, in the current exchange group processing method, it is usually based on the historical behavior data of the user to determine whether the user can join a certain exchange group, or through manual review to filter out users who are really interested in the theme of a certain exchange group, in order to ensure the quality of members of exchange groups of different theme types. For example, when a new user applies to join a certain exchange group, he needs to pass the manual review and approval before joining. This processing method has great limitations, and the administrator needs to review the application objects one by one when applying to join the current exchange group. In the case of a large number of application objects, it is easy to cause low processing efficiency of the exchange group. SUMMARY

[0004] Therefore, it is necessary to provide an exchange group processing method, device, computer equipment, computer readable storage medium and computer program product to effectively improve the processing efficiency of the exchange group and bring convenience to users.

[0005] In a first aspect, the present application provides an exchange group processing method. The method comprises: obtaining visual features and keyword features corresponding to a video submitted by a target object; determining a quality score of the video based on the visual features; determining a visual similarity between the visual features and preset visual features of an exchange group; determining a keyword similarity between the keyword features and preset keyword features of the exchange group; and determining admission information of the target object to join the exchange group based on the video quality score, the visual similarity and the keyword similarity.

[0006] In a second aspect, the present application provides a device for processing an exchange group. The device comprises: an obtaining module configured to obtain visual features and keyword features corresponding to a video submitted by a target object; a determining module configured to determine a quality score of the video based on the visual features; determine a visual similarity between the visual features and preset visual features of the exchange group; determine a keyword similarity between the keyword features and preset keyword features of the exchange group; and determine admission information of the target object to join the exchange group based on the quality score of the video, the visual similarity and the keyword similarity.

[0007] In an embodiment, the device further comprises: an extracting module configured to extract video content features of the video and use the video content features as the visual features; and extract keyword features corresponding to meta information of the video, wherein the meta information is information used to describe content and attributes of the video.

[0008] In an embodiment, the video content features comprise feature vectors; the device further comprises: an extracting module configured to extract key frames in the video; an adjusting module configured to adjust a size of the key frames to obtain adjusted key frames; and the extracting module is further configured to extract feature vectors of the adjusted key frames and use the feature vectors as the visual features.

[0009] In an embodiment, the extracting module is further configured to perform feature extraction on the adjusted key frames through a convolution layer of a video feature extraction model to obtain local feature vectors of the adjusted key frames.

[0010] In an embodiment, the meta information at least comprises title information and a description statement of the video; the device further comprises: a combining module configured to combine the title information and the description statement to obtain combined text information; and the determining module is further configured to determine a word vector of the combined text information and use the word vector as the keyword features corresponding to the meta information.

[0011] In an embodiment, the device further comprises: a processing module configured to perform classification processing on the visual features through a classification model to obtain the quality score of the video.

[0012] In an embodiment, the device further comprises: a calculating module configured to calculate an initial visual similarity between the visual features and the preset visual features; and a processing module configured to perform normalization processing on the initial visual similarity to obtain the visual similarity.

[0013] In one embodiment, the apparatus further includes: a calculation module, configured to calculate the cosine similarity between the keyword features and the preset word features; or to calculate the Euclidean distance between the keyword features and the preset word features; and a processing module, configured to normalize the cosine similarity or the Euclidean distance to obtain the keyword similarity.

[0014] In one embodiment, the determining module is further configured to determine a comprehensive score based on the weighting coefficient, the video quality score, the visual similarity, and the keyword similarity; and to determine the admission information for the target object to join the communication group based on the comprehensive score and risk control data.

[0015] In one embodiment, the weighting coefficients include a first weighting coefficient, a second weighting coefficient, and a third weighting coefficient; the determining module is further configured to determine, based on the characteristics of the communication group, the first weighting coefficient corresponding to the visual similarity and the second weighting coefficient corresponding to the keyword similarity; determine a video matching score based on the visual similarity, the keyword similarity, the first weighting coefficient, and the second weighting coefficient; determine a target quality score based on the video quality score and the third weighting coefficient; and determine the comprehensive score based on the video matching score and the target quality score.

[0016] In one embodiment, the determining module is further configured to determine a first product between the visual similarity and the first weight coefficient; determine a second product between the keyword similarity and the second weight coefficient; and use the sum of the first product and the second product as the video matching score.

[0017] In one embodiment, the apparatus further includes: a generation module, configured to generate admission information for rejecting the target object from joining the communication group when the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data meets the risk control conditions; and a determination module, configured to determine the admission information for the target object to join the communication group based on the comprehensive score when the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data does not meet the risk control conditions.

[0018] In one embodiment, the apparatus further includes: a comparison module, configured to compare the comprehensive score with a preset score to obtain a comparison result; a sending module, configured to send the application information of the target object corresponding to the comprehensive score to an administrator device when the comparison result indicates that the comprehensive score is greater than the preset score, so that the administrator device determines the admission information of the target object to join the communication group; a sorting module, configured to sort the comprehensive score with the comprehensive scores of other objects to obtain a sorted comprehensive score; and a determining module, further configured to determine the admission information of the target object to join the communication group from the sorted comprehensive scores.

[0019] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, performs the following steps: acquiring visual features and keyword features corresponding to a video submitted by a target object; determining a quality score for the video based on the visual features; determining the visual similarity between the visual features and preset visual features of a chat group; determining the keyword similarity between the keyword features and preset word features of the chat group; and determining the access information for the target object to join the chat group based on the video quality score, the visual similarity, and the keyword similarity.

[0020] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps: acquiring visual features and keyword features corresponding to a video submitted by a target object; determining a quality score for the video based on the visual features; determining the visual similarity between the visual features and preset visual features of a chat group; determining the keyword similarity between the keyword features and preset word features of the chat group; and determining the access information for the target object to join the chat group based on the video quality score, the visual similarity, and the keyword similarity.

[0021] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps: acquiring visual features and keyword features corresponding to a video submitted by a target object; determining a quality score for the video based on the visual features; determining the visual similarity between the visual features and preset visual features of a chat group; determining the keyword similarity between the keyword features and preset word features of the chat group; and determining access information for the target object to join the chat group based on the video quality score, the visual similarity, and the keyword similarity.

[0022] The aforementioned processing method, apparatus, computer equipment, storage medium, and computer program product for the chat group acquires visual and keyword features corresponding to the videos submitted by the target users; determines the video quality score based on the visual features; determines the visual similarity between the visual features and the chat group's preset visual features; determines the keyword similarity between the keyword features and the chat group's preset word features; and determines the target users' admission information for joining the chat group based on the video quality score, visual similarity, and keyword similarity. Since the visual and keyword features are automatically determined based on the videos submitted by the target users, the admission information for each target user to join the chat group can be quickly and accurately determined based on the visual features, keyword features, and the chat group's preset visual and word features. Even with a large number of applicants, the applications of a large number of users can be evaluated and processed in a short time, thereby effectively improving the processing efficiency of the chat group's admission information and bringing convenience to users. Attached Figure Description

[0023] Figure 1 This is a diagram illustrating the application environment of a communication group processing method in one embodiment.

[0024] Figure 2 This is a flowchart illustrating a method for processing a chat group in one embodiment;

[0025] Figure 3 This is a schematic diagram of the overall process of processing communication groups in a social application in one embodiment;

[0026] Figure 4 This is a flowchart illustrating the steps for obtaining the visual features and keyword features corresponding to the video submitted by the target object in one embodiment.

[0027] Figure 5 This is a flowchart illustrating the steps for extracting keyword features corresponding to metadata from a video in one embodiment.

[0028] Figure 6 This is a schematic diagram illustrating the extraction of visual features using a CNN in one embodiment.

[0029] Figure 7 This is a schematic diagram of visual feature extraction in one embodiment;

[0030] Figure 8 This is a flowchart illustrating the calculation of keyword similarity score and visual feature similarity score in one embodiment;

[0031] Figure 9 This is a structural block diagram of a communication group processing device in one embodiment;

[0032] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0034] It should be noted that in the following description, the terms "first, second, and third" are used only to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0035] The chat group processing method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. That is, terminal 102 can interact with a multimedia material publishing platform with chat group functionality, i.e., server 104. Server 104 obtains the visual features and keyword features corresponding to the video submitted by the target object through terminal 102. Server 104 determines the video's quality score based on the visual features and determines the visual similarity between the visual features and the preset visual features of the chat group; server 104 determines the keyword similarity between the keyword features and the preset word features of the chat group; server 104 determines the access information for the target object to join the chat group based on the video quality score, visual similarity, and keyword similarity. For example, when server 104 determines that the target object's access information for joining the chat group is allowed based on the video quality score, visual similarity, and keyword similarity, server 104 can return the access information indicating permission to join the chat group to terminal 102 (i.e., the terminal device used by the target object).

[0036] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, IoT device, or portable wearable device. IoT devices can include smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc.

[0037] Server 104 can be an independent physical server or a service node in a blockchain system. The service nodes in the blockchain system form a peer-to-peer (Peer To Peer) network. The Peer To Peer protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP).

[0038] In addition, server 104 can also be a server cluster consisting of multiple physical servers, which can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0039] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.

[0040] In one embodiment, such as Figure 2 As shown, a method for processing chat groups is provided. This method can be executed by the server or the terminal alone, or by both the server and the terminal. This method can be applied to... Figure 1 Taking the server in the example, the following steps are included:

[0041] Step 202: Obtain the visual features and keyword features corresponding to the video submitted by the target object.

[0042] The target group refers to the applicants who have applied to join the discussion group, that is, one or more applicants who have applied to join the discussion group. For example, if user A and user B both submit applications to join interest group A (a certain type of discussion group), then the target group includes user A and user B.

[0043] The term "video" refers to a video submitted by the target entity. For example, the video submitted by the target entity in this application may include local videos, short videos captured from the internet, etc. Furthermore, the methods by which the target entity submits videos in this application include, but are not limited to: the target entity may submit a link to a short video, or the target entity may select a local video to upload and use the uploaded local video as the submitted video.

[0044] Visual features refer to the video content features corresponding to the video submitted by the target object. For example, in this application, the visual features can be the video content features A obtained by processing the video submitted by the target object through a video feature extraction model, and then using the output video content features A as the visual features corresponding to the video submitted by the target object.

[0045] Keyword features refer to the keyword features corresponding to the metadata of the video submitted by the target object. For example, in this application, the keyword features can be obtained by processing the metadata of the video submitted by the target object based on a keyword extraction algorithm, and then using the processed keyword feature B as the keyword feature corresponding to the video submitted by the target object. Meta-information refers to descriptive information used to describe the content and attributes of the video submitted by the target object. For example, the metadata in this application may include titles, descriptive statements, subtitles, tags, etc.

[0046] Specifically, the method for processing communication groups provided in this application can be widely applied to various personalized material creation and publishing fields such as social networking, games, and film and television production. That is, the devices used by different users (operation objects) can interact with the multimedia information platform (or application). When a user (operation object) wants to apply to join a communication group (such as interest group A), the user can open the multimedia application (Application, APP) on the terminal by triggering an operation, and enter the main page of the multimedia application by selecting an operation. That is, the user can log in to the multimedia application (such as video material publishing application, social application) by triggering an operation. Furthermore, the user can initiate a request to join the communication group by triggering an operation on the main page displayed by the multimedia application. For example, on the main page of a social application (video account application) displayed on the terminal, user A can submit a video by clicking the "Submit Material" icon. The terminal then responds to this "Submit Material" request triggered by user A in the social application by sending the request to the application's backend server. The backend server can then obtain the video submitted by user A, extract its content features, and use these features as visual features. Simultaneously, the backend server can also extract keyword features corresponding to the metadata of the submitted video and use these keyword features as the keyword features of the submitted video.

[0047] For example, let's take the scenario of handling a communication group in a social application as an example. When user A (the target) wants to apply to join a certain interest group A (communication group), user A can open a specific social application on the terminal through a trigger operation. That is, the user can log in to the social application through a trigger operation. Furthermore, the user can initiate a request to submit multimedia materials (or a request to join the communication group) through a trigger operation on the page displayed in the social application. For example, on a social application (video account application) page displayed on a terminal, user A can view the specific content and related functional information on the page. Suppose user A clicks the "Submit Video" icon or the "Apply to Join Discussion Group" icon on the page and submits the link to video A. In response to the "Submit Video" or "Apply to Join Discussion Group" request triggered by user A in the social application, the terminal sends the "Submit Video" or "Apply to Join Discussion Group" request triggered by user A to the backend server of the social application. The backend server of the social application can then obtain the target object, namely the link to video A submitted by user A, and extract the video content features of video A through a video feature extraction model, and use the extracted video content features as the visual features of video A. At the same time, the backend server can also extract the title information and keyword features in the description of video A, and use the extracted keyword features as the keyword features of video A.

[0048] Step 204: Determine the quality score of the video based on visual features.

[0049] The quality score refers to an indicator used to evaluate the content quality of a video. For example, the quality score in this application may include different score ranges from 0 to 100. When the quality score is 0, it means that the video quality is the worst, and when the quality score is 100, it means that the video quality is the best.

[0050] Specifically, after the backend server obtains the visual and keyword features corresponding to the video submitted by the target object, it can determine the video's quality score based on the visual features. For example, the backend server can input the visual features corresponding to the video into a pre-trained neural network (model). After processing by the neural network, the output value of the neural network is the quality score. This neural network can be understood as a classification model that can classify the video's visual features into different score ranges from 0 to 100, where 0 represents the worst video quality and 100 represents the best video quality.

[0051] It is understood that this application used a large amount of pre-collected and pre-scored short video data when training the neural network (i.e., the classification model). This data includes short videos of various types and qualities, which can help the neural network learn the relationship between video quality and visual features.

[0052] Step 206: Determine the visual similarity between the visual features and the preset visual features of the communication group.

[0053] Among them, the discussion group refers to the discussion group created for different topics. For example, if user A likes photography, then user A can apply to join a photography interest group (i.e., a photography discussion group).

[0054] Preset visual features refer to the pre-defined video content features corresponding to a specific chat group. For example, when a group administrator creates or manages an interest group (chat group), they upload short videos representing the group's theme and a descriptive text. The backend server processes this information. For instance, the backend server uses the TF-IDF algorithm to extract keyword features from the descriptive text (as preset keyword features), and simultaneously uses a CNN algorithm to extract keyframes from the short videos uploaded by the administrator and extract visual features (as preset visual features). These features are stored as preset features for later use. Therefore, when an applicant applies to join the interest group, they also need to upload or submit a short video that they like or are interested in, so that the backend server can determine the admission information for different applicants to join different types of chat groups based on the visual and keyword features corresponding to the video submitted by the applicant.

[0055] Visual similarity refers to the similarity between the visual features of a video submitted by a target object and the preset visual features of a communication group. For example, the calculation methods of visual similarity in this application include, but are not limited to: calculating the cosine similarity between the visual features of the video submitted by the target object and the preset visual features of the communication group; or calculating the Euclidean distance between the visual features of the video submitted by the target object and the preset visual features of the communication group.

[0056] Step 208: Determine the keyword similarity between the keyword features and the preset word features of the communication group.

[0057] Among them, the preset word features refer to the preset keyword features corresponding to a specific communication group. For example, when a group administrator creates or manages an interest group (communication group), he will upload some short videos and a description representing the theme of the interest group. The backend server will process this information. For example, the backend server uses the TF-IDF algorithm to extract the keyword features of the description and uses the keyword features of the description as the preset word features corresponding to the interest group (communication group).

[0058] Keyword similarity refers to the similarity between the keyword features of the video submitted by the target object and the preset word features of the communication group. For example, the calculation method of keyword similarity in this application includes, but is not limited to: calculating the cosine similarity between the keyword features of the video submitted by the target object and the preset word features of the communication group; or calculating the Euclidean distance between the keyword features of the video submitted by the target object and the preset word features of the communication group.

[0059] Specifically, after the backend server determines the quality score of the video based on visual features, it can process the visual features and keyword features corresponding to the video separately. That is, the backend server can calculate the initial visual similarity between the visual features and the preset visual features of the chat group, and normalize the initial visual similarity to obtain the visual similarity. At the same time, the backend server can calculate the similarity between the keyword features and the preset word features of the chat group, and normalize the calculated similarity to obtain the keyword similarity.

[0060] For example, let's take the processing scenario of a chat group in a social application as an example. Suppose the visual features corresponding to the video A submitted by the target object, as video content features S, and the keyword features, as A1, obtained by the backend server. The backend server can calculate the Euclidean distance D1 between the visual features (video content features S) and the preset visual features S0 of chat group A. After normalizing the calculated Euclidean distance D1, a normalized similarity d1 is obtained, which is then used as the visual similarity. Furthermore, the backend server can calculate the Euclidean distance D2 between the keyword features A1 and the preset word features C0 of chat group A. After normalizing the calculated Euclidean distance D2, a normalized similarity d2 is obtained, which is then used as the keyword similarity.

[0061] Step 210: Based on video quality score, visual similarity, and keyword similarity, determine the admission information for the target object to join the communication group.

[0062] The access information refers to the access information regarding whether the target object can join the communication group. For example, the access information in this application may include information on whether joining is allowed or denied.

[0063] Specifically, after the backend server determines the keyword similarity between the keyword features and the preset word features of the chat group, the backend server can determine the admission information for the target object to join the chat group based on video quality score, visual similarity, and keyword similarity. For example, the backend server can calculate a comprehensive score based on weight coefficient, video quality score, visual similarity, and keyword similarity. Furthermore, the backend server can comprehensively determine the admission information for the target object to join the chat group based on risk control data and the calculated comprehensive score.

[0064] For example, such as Figure 3 The diagram shown illustrates the overall workflow for handling communication groups in social applications. Figure 3 As shown, the backend server can calculate a video matching score (i.e., matching score) based on visual similarity and keyword similarity. Simultaneously, the backend server can calculate a target quality score (i.e., quality score) based on the video quality score and weighting coefficients. The backend server can then comprehensively determine whether a target individual is eligible to join the chat group based on risk control information, the calculated video matching score, and the target quality score. For example, if the relevance of the target individual's historical application success rate, historical application frequency, and historical application topics in the risk control information (i.e., risk control data) meets the risk control conditions, the backend server can directly generate admission information rejecting the target individual from joining the chat group. If the relevance of the target individual's historical application success rate, historical application frequency, and historical application topics in the risk control information (i.e., risk control data) does not meet the risk control conditions, the backend server can further comprehensively determine whether the target individual is eligible to join the chat group based on the calculated video matching score and the target quality score.

[0065] In this embodiment, visual and keyword features corresponding to the videos submitted by the target object are obtained; the quality score of the video is determined based on the visual features; the visual similarity between the visual features and the preset visual features of the chat group is determined; the keyword similarity between the keyword features and the preset word features of the chat group is determined; and the admission information for the target object to join the chat group is determined based on the video quality score, visual similarity, and keyword similarity. Since the visual and keyword features are automatically determined based on the videos submitted by the target object, the admission information for each target object to join the chat group can be quickly and accurately determined based on the visual features, keyword features, preset visual features of the chat group, and preset word features. Even with a large number of applicants, the applications of a large number of users can be evaluated and processed in a short time, thereby effectively improving the processing efficiency of the admission information for the chat group and bringing convenience to users.

[0066] In one embodiment, such as Figure 4 As shown, the steps for obtaining the visual features and keyword features corresponding to the video submitted by the target object include:

[0067] Step 402: Extract video content features from the video and use these video content features as visual features;

[0068] Step 404: Extract keyword features corresponding to the video's metadata; metadata is information used to describe the video's content and attributes.

[0069] The metadata in this application embodiment may include: title, description, subtitle, tag (category tag), etc.

[0070] Specifically, let's take the processing scenario of an interest-based exchange group in a social application as an example. When user A (the target) wants to apply to join an interest group A (exchange group) on a certain topic about food, user A can open a specific social application A on the terminal through a triggering action. User A can initiate a request to join interest group A on the page displayed by social application A. For example, suppose user A clicks the "Submit Video" icon or the "Apply to Join Exchange Group" icon on the page displayed by social application A and submits the link to video A. In response to the "Submit Video" request or "Apply to Join Exchange Group" request triggered by user A in social application A, the terminal sends the "Submit Video" request or "Apply to Join Exchange Group" request triggered by user A to the backend server of social application A. The backend server of social application can then obtain video A based on the link to video A. The backend server can extract the video content feature S of video A and use the video content feature S as a visual feature. At the same time, the backend server can extract the keyword feature S1 corresponding to the metadata of video A and use the keyword feature S1 corresponding to the metadata as the keyword feature corresponding to video A. This allows for the evaluation of interest and research level of any user-submitted video by having users actively submit short video content. Since this process is initiated by the user and can be processed by the backend server at any time, it can effectively improve the processing efficiency of access information for different communication groups and bring convenience to users.

[0071] In one embodiment, the video content features include feature vectors; the step of extracting the video content features and using the video content features as visual features includes:

[0072] Extract keyframes from the video;

[0073] Adjust the size of the keyframe to obtain the adjusted keyframe;

[0074] Extract the feature vectors of the adjusted keyframes and use the feature vectors as visual features.

[0075] In this application, keyframes are static images in a video that represent a certain moment in time. Keyframes can be selected using various methods. The method of extracting keyframes in this application may include: selecting one frame at fixed time intervals s, for example, extracting one frame from the video as a keyframe every s = 3 seconds.

[0076] Specifically, let's take the processing scenario of an interest-based discussion group in a social application as an example. Suppose user A triggers a "Request to join discussion group" request in social application A and submits a link to video A. In response to user A's "Request to join discussion group" request, the terminal sends the request to the backend server of social application A. When the backend server receives the request and the link to video A, it can retrieve video A based on the link and extract keyframes from video A at fixed time intervals. Furthermore, the backend server can adjust the size of the keyframes to obtain adjusted keyframes, extract the feature vectors of the adjusted keyframes, and use the extracted feature vectors as the visual features corresponding to video A. It can be understood that the reason for adjusting the size of the keyframes in this embodiment is to adjust the extracted keyframes to a fixed size to meet the input requirements of the model. This enables a series of automated processing steps, including video content and topic extraction, similarity determination, and video content quality scoring, to evaluate a large number of user applications in a short time. This reduces the subjectivity and inconsistency that can easily occur with manual review, effectively improving the objectivity and efficiency of the approval results for each target to join the communication group, and bringing convenience to users.

[0077] In one embodiment, the step of extracting the feature vector of the adjusted keyframe includes:

[0078] The adjusted keyframes are processed by using the convolutional layer of the video feature extraction model to extract features, resulting in local feature vectors for the adjusted keyframes.

[0079] Among them, the video feature extraction model refers to the model used to extract video content features. For example, the video feature extraction model used in this application can be a convolutional neural network (CNN).

[0080] Specifically, let's take the processing scenario of interest-based discussion groups in a social application as an example. Suppose user A triggers a "Request to join discussion group" request in social application A and submits a link to video A. When the backend server of social application A receives the "Request to join discussion group" request and the link to video A initiated by user A, the backend server can obtain video A based on the link and extract keyframes from video A at fixed time intervals. Furthermore, the backend server can adjust the size of the keyframes to obtain adjusted keyframes, and input the adjusted keyframes into a pre-trained video feature extraction model for processing. That is, the convolutional layer of the video feature extraction model extracts features from the adjusted keyframes. After processing by the convolutional layer of the video feature extraction model, the local feature vector S of the adjusted keyframes is output, and the output local feature vector S is used as the visual feature corresponding to video A. The main function of the convolutional layer is to extract local features of the image, such as edges, corners, and textures. A convolutional layer contains multiple convolutional kernels, each capturing a specific feature in the image. The kernels perform a sliding window operation on the input image, and the dot product between the kernel and a local region of the input image is calculated to obtain the convolutional feature map. It is understood that the reason for adjusting the keyframe size in this embodiment is to adjust the extracted keyframes to a fixed size to meet the model's input requirements. This allows for effective improvement of the processing efficiency of the chat group by using a video feature extraction model to extract visual features from user-submitted short videos, bringing convenience to users.

[0081] In one embodiment, such as Figure 5 As shown, the metadata includes at least the video's title and description; the steps for extracting keyword features corresponding to the video's metadata include:

[0082] Step 502: Combine the title information and the description statement to obtain combined text information;

[0083] Step 504: Determine the word vectors of the combined text information and use the word vectors as the keyword features corresponding to the meta-information.

[0084] The title information refers to the title of the video submitted by the target object, while the description statement refers to the statement used to describe the video submitted by the target object. For example, suppose there is a short video A with the title "Amazing underwater diving in Southeast Asia" and the description statement (description information) is "Diving in Southeast Asia is an unforgettable experience. The underwater world is full of colorful fishes and corals".

[0085] Specifically, let's take short video A submitted by the target user as an example. Assume user A triggers a "Request to join a discussion group" request in social application A and submits a link to short video A. When the backend server of social application A receives the "Request to join a discussion group" request and the link to video A, the backend server can obtain short video A and its title S1 ("Amazing underwater diving in Southeast Asia") and description S2 ("Diving in Southeast Asia is an unforgettable experience. The underwater world is full of colorful fishes and corals") based on the link. Furthermore, the backend server can combine the title S1 and description S2 of short video A to obtain combined text information S1+S2, and process this text information (i.e., combined text information S1+S2) using the TF-IDF algorithm to obtain the word vector T of the processed text information (i.e., combined text information S1+S2). The word vector T is then used as the keyword feature corresponding to the metadata of short video A. For example, assuming the vocabulary corresponding to the combined text information S1+S2 contains words such as "amazing", "underwater", "diving", and "Asia", then the output word vector T might be as follows:

[0086]

[0087] This word vector represents the importance of words such as "amazing," "underwater," "diving," and "Asia" in the text. The backend server can use this word vector as a keyword feature for short video A, which can then be used for subsequent feature fusion and similarity calculation. This enables a series of automated processes, including video content and topic extraction, similarity determination, and video content quality scoring, to evaluate a large number of user applications in a short time. This reduces the subjectivity and inconsistencies that can easily occur with manual review, effectively improving the objectivity and efficiency of approving applications for joining the discussion group and bringing convenience to users.

[0088] In one embodiment, the step of determining a video quality score based on visual features includes:

[0089] The visual features are classified using a classification model to obtain a quality score for the video.

[0090] The classification model refers to a model used to classify videos (visual features) into different score ranges from 0 to 100. For example, the classification model in this application can be a pre-trained neural network model.

[0091] Specifically, after the backend server obtains the visual and keyword features corresponding to video A submitted by the target object, it can input the visual features of video A into a pre-trained neural network. After processing by the neural network, the output value is the quality score of video A. This neural network can be understood as a classification model that can categorize the visual features of a video into different score ranges from 0 to 100, where 0 represents the worst video quality and 100 represents the best. Therefore, by using a classification model to evaluate the quality of user-submitted short video content, reflecting the user's research and enthusiasm for the topic, the accuracy of the admission information for different users to join specific discussion groups can be effectively improved. This ensures that the short videos submitted by users when applying to join different interest groups (discussion groups) are authentic and valid, preventing arbitrary video selection or malicious mass applications.

[0092] In one embodiment, the step of determining the visual similarity between visual features and preset visual features of a communication group includes:

[0093] Calculate the initial visual similarity between visual features and preset visual features;

[0094] The initial visual similarity is normalized to obtain the visual similarity.

[0095] The initial visual similarity refers to the original similarity obtained by calculation. For example, in the embodiments of this application, Euclidean distance can be used to calculate the initial visual similarity between visual features and preset visual features.

[0096] Specifically, let's take the processing scenario of a chat group in a social application as an example. Assume the visual features corresponding to the video A submitted by the target object, as video content feature A, and the keyword features, as A1, obtained by the backend server. The backend server can calculate the Euclidean distance D1 between the visual feature (video content feature A) and the preset visual feature B of chat group A. The calculated Euclidean distance D1 is then normalized to a score range of 0-100, yielding the normalized similarity d1. This normalized similarity d1 is then used as the visual similarity. Therefore, by extracting video content features and thematic information features from user-submitted selected or created short videos, more accurate data is provided for subsequent similarity judgment and quality assessment. Even with a large number of applicants, applications from a large number of users can be evaluated and processed in a short time, effectively improving the processing efficiency of access information for different chat groups and bringing convenience to users.

[0097] The Euclidean distance between vector A (i.e., video content feature A) and vector B (i.e., preset visual feature B) can be calculated using the following formula (1):

[0098] (1)

[0099] Where n represents the dimension of the vector, A i and B i Let A and B represent the i-th values ​​of vectors A and B, respectively.

[0100] In one embodiment, the step of determining the keyword similarity between keyword features and preset word features of a communication group includes:

[0101] Calculate the cosine similarity between keyword features and preset word features; or

[0102] Calculate the Euclidean distance between keyword features and preset word features;

[0103] The cosine similarity or Euclidean distance is normalized to obtain the keyword similarity.

[0104] Specifically, let's take the processing scenario of a communication group in a social application as an example. Assume the visual features corresponding to the video A submitted by the target object, obtained by the backend server, are video content features A, and the keyword features are A1. The backend server can calculate the cosine similarity D0 between keyword feature A1 and the preset word feature S0; or the backend server can calculate the Euclidean distance D1 between keyword feature A1 and the preset word feature S0, and normalize the calculated cosine similarity D0 or Euclidean distance D1, that is, normalize the cosine similarity D0 or Euclidean distance D1 to a score range of 0-100, thus obtaining the normalized similarity d0 or d1, which is then used as the keyword similarity. Therefore, by extracting video content features and thematic information features from the selected or created short videos submitted by users, more accurate data is provided for subsequent similarity judgment and quality assessment. Even with a large number of applicants, a large number of user applications can be evaluated and processed in a short time, effectively improving the processing efficiency of access information for different communication groups and bringing convenience to users.

[0105] In one embodiment, the step of determining the admission information for a target object to join a communication group based on video quality score, visual similarity, and keyword similarity includes:

[0106] The comprehensive score is determined based on weighting coefficients, video quality scores, visual similarity, and keyword similarity.

[0107] Based on comprehensive scores and risk control data, the admission information for target individuals to join the communication group is determined.

[0108] The weighting coefficients in this application may include weighting coefficients corresponding to different evaluation indicators (parameters). For example, the weighting coefficient corresponding to the video quality score is the third weighting coefficient, the weighting coefficient corresponding to visual similarity is the first weighting coefficient, and the weighting coefficient corresponding to keyword similarity is the second weighting coefficient. The sum of the third, first, and second weighting coefficients corresponding to the video quality score, visual similarity, and keyword similarity is 1. It can be understood that this application can dynamically determine the third, first, and second weighting coefficients corresponding to the video quality score, visual similarity, and keyword similarity based on different weighting allocation strategies.

[0109] Specifically, such as Figure 3As shown, the backend server can determine a comprehensive score based on weighted coefficients, video quality scores, visual similarity, and keyword similarity. Furthermore, the backend server can determine the admission criteria for a target individual to join the chat group based on the comprehensive score and risk control data. For example, when the relevance of the target individual's historical application success rate, historical application frequency, and historical application topics in the risk control data meets the risk control conditions, the backend server can directly generate admission criteria to reject the target individual from joining the chat group. When the relevance of the target individual's historical application success rate, historical application frequency, and historical application topics in the risk control data does not meet the risk control conditions, the backend server can further determine whether the target individual can join the chat group based on the calculated comprehensive score. This allows for a comprehensive score based on the similarity between short video content and the theme of the interest group, as well as the quality of the short video content. This score is then used as the basis for interest group admission, effectively improving the processing efficiency of admission information for different communication groups. It also includes necessary risk control measures, using various strategies and technologies to ensure that the short videos submitted by users when applying to join interest groups are authentic and valid, preventing arbitrary video selection or malicious mass applications.

[0110] In one embodiment, the weighting coefficients include a first weighting coefficient, a second weighting coefficient, and a third weighting coefficient; the method further includes:

[0111] Based on the characteristics of the communication group, the first weight coefficient corresponding to visual similarity and the second weight coefficient corresponding to keyword similarity are determined.

[0112] The steps for determining the comprehensive score based on weighted coefficients, video quality scores, visual similarity, and keyword similarity include:

[0113] The video matching score is determined based on visual similarity, keyword similarity, first weight coefficient, and second weight coefficient.

[0114] The target quality score is determined based on the video quality score and the third weighting coefficient.

[0115] The overall score is determined based on the video matching score and the target quality score.

[0116] The characteristics of a chat group refer to information reflecting the preferences of that chat group. For example, if the characteristic of chat group A is primarily video content (or visual experience), then for chat group A, a larger proportion of the first weight coefficient corresponding to visual similarity can be assigned, meaning the backend server can determine that the first weight coefficient corresponding to visual similarity is greater than the second weight coefficient corresponding to keyword similarity. Similarly, if the characteristic of chat group B is primarily video topic relevance, then for chat group B, a larger proportion of the second weight coefficient corresponding to keyword similarity can be assigned, meaning the backend server can determine that the second weight coefficient corresponding to keyword similarity is greater than the first weight coefficient corresponding to visual similarity.

[0117] Specifically, such as Figure 3 As shown, the backend server can dynamically determine the first weight coefficient corresponding to visual similarity and the second weight coefficient corresponding to keyword similarity based on the characteristics of the chat group. Furthermore, the backend server can determine, based on visual similarity, keyword similarity, the first weight coefficient, and the second weight coefficient, as follows: Figure 3 The matching score shown is the video matching score, which is determined based on the video quality score and the third weighting coefficient. Figure 3 The quality score shown is the target quality score. The backend server can determine a comprehensive score based on the video matching score and the target quality score, so that subsequent admission information for the target user to join the chat group can be determined based on the comprehensive score. This allows for the evaluation and processing of a large number of user applications in a short time, even when the number of applicants is large, thus effectively improving the processing efficiency of admission information for different chat groups and bringing convenience to users.

[0118] In one embodiment, the step of determining the video matching score based on visual similarity, keyword similarity, a first weighting coefficient, and a second weighting coefficient includes:

[0119] Determine the first product between visual similarity and the first weighting coefficient;

[0120] Determine the second product between keyword similarity and the second weighting coefficient;

[0121] The sum of the first and second products is used as the video matching score.

[0122] Specifically, such as Figure 3 As shown in the diagram, the backend server can dynamically determine the first weight coefficient A1 corresponding to visual similarity d1 and the second weight coefficient A2 corresponding to keyword similarity d2 based on the characteristics of the communication group A. Furthermore, the backend server can determine the following based on visual similarity d1, keyword similarity d2, the first weight coefficient A1, and the second weight coefficient A2: Figure 3The matching score shown is the video matching score S. The backend server calculates the first product between visual similarity d1 and the first weight coefficient A1 as d1×A1, and the second product between keyword similarity d2 and the second weight coefficient A2 as d2×A2. The sum of the first product d1×A1 and the second product d2×A2 is taken as the video matching score S = (d1×A1) + (d2×A2). This allows for a more accurate assessment of different users' attention to the interest group's theme by extracting content features and theme information features from user-submitted selected or created short videos. By comparing the similarity between the extracted short video content features and the interest group's preset short video content features, and the similarity between the theme information features and the interest group's preset theme features, the server can quickly and accurately determine the admission information for each target user to join the group, effectively improving the accuracy of admission information for specific target users.

[0123] In one embodiment, the step of determining the access information for a target to join a communication group based on a comprehensive score and risk control data includes:

[0124] When the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data meets the risk control conditions, access information is generated to reject the target object from joining the communication group.

[0125] When the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data does not meet the risk control conditions, the admission information for the target to join the communication group is determined based on the comprehensive score.

[0126] The risk control data includes at least the applicant's historical application success rate, (historical) application frequency, and the relevance of historical application topics.

[0127] Specifically, after the backend server determines the comprehensive score, it can first check the risk control data through the anti-cheating and anti-malicious application risk control module. Specifically, if the relevance of the target user's historical application success rate, historical application frequency, and historical application topics in the risk control data meets the risk control conditions, the backend server can generate admission information to reject the target user from joining the exchange group. Alternatively, if the relevance of the target user's historical application success rate, historical application frequency, and historical application topics in the risk control data does not meet the risk control conditions, the backend server can further determine whether the target user can join the exchange group based on the comprehensive score. This allows for a comprehensive score based on the similarity between short video content and the interest group's theme, as well as the quality of the short video content, to assess different users' interest and understanding of specific interest groups. This comprehensive score serves as the basis for interest group admission, effectively improving the processing efficiency of admission information for different exchange groups. It also includes necessary risk control measures, ensuring that the short videos submitted by users when applying to join interest groups are authentic and valid, preventing arbitrary video selection or malicious mass applications.

[0128] In one embodiment, the step of determining the admission information for a target object to join the chat group based on a comprehensive score includes:

[0129] The overall score is compared with the preset score to obtain the comparison result; when the comparison result indicates that the overall score is greater than the preset score, the application information of the target object corresponding to the overall score is sent to the administrator device so that the administrator device can determine the access information for the target object to join the communication group; or

[0130] The overall score is sorted with the overall scores of other objects to obtain the sorted overall score; the admission information for the target object to join the communication group is determined from the sorted overall score.

[0131] The preset score refers to a pre-set threshold; for example, the preset score in this application can be 80.

[0132] Specifically, when the relevance of the target object's historical application success rate, historical application frequency, and historical application topics in the risk control data does not meet the risk control conditions, the backend server can further determine whether the target object can join the communication group based on the comprehensive score. For example, the backend server can compare the comprehensive score with a preset score to obtain a comparison result; when the comparison result indicates that the comprehensive score is greater than the preset score, the backend server can send the application information of the target object corresponding to the comprehensive score to the administrator device, so that the administrator device can determine the target object's access information to join the communication group.

[0133] Alternatively, the backend server can sort the overall score with the overall scores of other objects to obtain a sorted overall score, and then determine the admission information for the target object to join the communication group from the sorted overall scores. This effectively improves the processing efficiency of admission information for different communication groups, bringing convenience to users.

[0134] This application also provides an application scenario in which the above-mentioned chat group processing method is applied. Specifically, the application of the chat group processing method in this scenario is as follows:

[0135] During user interaction with a multimedia information (or image editing) platform, the aforementioned method for handling chat groups can be used. When a user (the object of operation) wants to apply to join a chat group (e.g., interest group A), the user can trigger an action to open a multimedia application on their terminal. That is, the user can log in to a multimedia application (e.g., a video content publishing application, a social application) through a trigger action. Furthermore, the user can initiate a request to join the chat group through a trigger action on the page displayed in the multimedia application. When the backend server of the social application receives the user's request to join the chat group, the backend server can obtain the video A submitted by the target object (i.e., the object of operation). The server can extract video content features from the submitted video A and use these features as visual features. Simultaneously, the backend server can extract keyword features corresponding to the metadata of the submitted video A and use these features as keyword features. Furthermore, the backend server can determine the video's quality score based on the visual features, and determine the visual similarity between the visual features and the preset visual features of the chat group, as well as the keyword similarity between the keyword features and the preset word features of the chat group. Based on the video quality score, visual similarity, and keyword similarity, the backend server determines the admission information for the target user to join the chat group. Therefore, when different users interact with the multimedia information platform, the chat group processing method provided in this application, through automated short video content and topic extraction, similarity determination, and content quality scoring, can evaluate a large number of user applications in a short time. This reduces the subjectivity and inconsistency that easily occur in traditional manual review methods, thereby effectively improving the objectivity and efficiency of the approval results for different chat groups.

[0136] The method provided in this application can be applied to various scenarios involving the processing of chat groups. The following example illustrates the chat group processing method provided in this application.

[0137] Video features, in particular, are a set of numerical or vector values ​​extracted from video content that describe the video's attributes and information. Video features are typically used to represent the visual, audio, and semantic information of a video, facilitating computer processing and analysis. By extracting video features, complex video content can be transformed into a simplified representation, enabling tasks such as video classification, retrieval, and similarity matching.

[0138] Meta-information: Information used to describe data, providing detailed information about the data and helping to understand its source, structure, characteristics, etc. Meta-information can be used for data management, searching, organization, and understanding. In short videos, meta-information may include titles, descriptions, tags (category tags), captions, etc., used to describe the video's content and attributes.

[0139] Normalization is a data preprocessing technique primarily used to eliminate the influence of differences in data dimensions and scales, transforming data to a uniform standard or range. The purpose of normalization is to make different features or variables comparable, facilitating subsequent calculations and analysis. In machine learning and data mining, normalization is a common data preprocessing method that can improve model convergence speed and performance.

[0140] Traditional technical solutions include:

[0141] 1. User verification system in community forums: To ensure the quality of the forum, new users need to undergo manual verification before joining. The verification is usually based on the user's application form, including information such as self-introduction, understanding of the forum topic, and enthusiasm.

[0142] 2. Interest Group Recommendation Systems on Social Media Platforms: Some social media platforms recommend relevant interest groups based on users' interests and behavioral history. This method can automatically recommend interest groups to users, improving the efficiency of the recommendation process.

[0143] In summary, while manual review systems can filter out users genuinely interested in forum topics, they are inefficient and rely on human review, potentially leading to subjectivity and inconsistencies. Social media platform interest group recommendation systems typically rely on users' historical behavioral data, which may not accurately reflect users' interest and enthusiasm for new topics, and users are passively receiving recommendations.

[0144] The technical solution provided in this application addresses the problems of traditional technologies through the following two points:

[0145] 1. Improved efficiency and objectivity: Through automated modules such as short video content and theme extraction, similarity determination, and content quality scoring, this solution can evaluate a large number of user applications in a short time, reducing the impact of subjectivity and inconsistency, and improving the objectivity and efficiency of the approval results.

[0146] 2. New Topics and Initiative: By allowing users to actively submit short video content, the system can assess the interest and research level of any topic proposed by the user. This process is initiated by the user and can be processed at any time by the system provided in this application.

[0147] The technical solution provided in this application utilizes the relevance of short video content to assess users' research and enthusiasm for interest group topics, thereby achieving more efficient and accurate interest group admission approval. By extracting content and themes, determining similarity, and evaluating the quality of user-submitted short videos, this technical solution can provide more suitable member recommendations for interest groups, thereby promoting communication and discussion within the groups and improving user satisfaction. Simultaneously, through a risk control module to prevent cheating and malicious applications, this technical solution can ensure the quality of interest group admissions and prevent malicious behavior from affecting the normal operation of interest groups.

[0148] The key technical aspects of this application include whether the content of short videos matches the theme of interest groups and the scoring mechanism for whether users are suitable to join interest groups. Specifically, this application proposes an interest group admission approval system based on the relevance of short video content, which consists of the following modules:

[0149] 1. Short Video Content and Theme Extraction Module: This module is used to extract content and theme information from user-submitted selected or created short videos for subsequent similarity judgment and quality assessment.

[0150] 2. Short video content and topic similarity determination module: This module compares the similarity between extracted short video content and topics and interest group topics to assess the user's level of interest in the topic.

[0151] 3. Short video content quality scoring module: Used to evaluate the quality of user-submitted short video content to reflect the user's level of research and enthusiasm for the topic.

[0152] 4. User Admission Comprehensive Evaluation Module: Based on the similarity between short video content and the theme, as well as the content quality score, a comprehensive score is given to the user's interest and understanding, serving as the basis for admission to interest groups. It also includes necessary risk control measures, using various strategies and technical means to ensure that the short videos submitted by users when applying to join interest groups are authentic and valid, preventing arbitrary video selection or malicious mass applications.

[0153] On the technical side, 1.1 Overall Process

[0154] like Figure 3 The diagram shows the overall process of this application mechanism. In the interest group access approval system provided in this application, the process mainly includes the following steps:

[0155] First, when creating or managing an interest group, the group administrator uploads short videos representing the group's theme and a descriptive text. The system provided in this application processes this information. Specifically, it uses the TF-IDF algorithm to extract keyword features from the descriptive text and a CNN algorithm to extract keyframes and visual features from the uploaded short videos. These features are stored for later use. When an applicant applies to join the interest group, they also need to upload their own short video. The system provided in this application performs the same feature extraction processing on the applicant's uploaded short video and its metadata to obtain the short video's keyword and visual features.

[0156] Next, the system provided in this application compares the keyword features of the applicant's short video with the keyword features of the interest group topics set by the administrator, and also compares the video features of both. A matching score is obtained by calculating the Euclidean distance to measure the similarity between the applicant's short video content and the interest group topics. Simultaneously, the system provides the system that inputs the visual features of the applicant's short video into a pre-trained neural network model for quality scoring. This neural network model is a multi-classification model that can classify short videos into different quality score ranges.

[0157] Finally, the system provided in this application will comprehensively evaluate applicants based on their matching score, quality score, and risk control information (such as historical application success rate, application frequency, and relevance between historical application topics) to obtain a comprehensive score. Based on the comprehensive score, the system will rank applicants, and those with higher scores will be allowed to join interest groups.

[0158] Through the above steps, the interest group admission approval system provided in this application can effectively assess applicants' interest in and understanding of the interest group's topic, ensuring the quality of admission and improving the effectiveness of communication and discussion within the interest group. To accomplish these steps, this application proposes four modules: a short video content and topic extraction module, a short video content and topic similarity determination module, a short video content quality scoring module, and a comprehensive user admission evaluation module. The following sections will detail the technologies and usage methods involved in each module:

[0159] 1.2 Short Video Content and Theme Extraction Module

[0160] 1.2.1 CNN for extracting visual features

[0161] like Figure 6 The diagram illustrates how a CNN extracts visual features. A Convolutional Neural Network (CNN) is a deep learning model primarily used for processing image and video data. The basic principle of a CNN is to automatically extract features from the input data using structures such as convolutional layers, pooling layers, and fully connected layers, and then perform classification or regression tasks. During training, CNNs continuously update the weight parameters of the convolutional kernels and fully connected layers using optimization methods such as backpropagation and gradient descent to improve the model's predictive performance.

[0162] To more clearly illustrate how visual features are extracted in this application, an example of convolution is given here, such as... Figure 7 The image shown is a schematic diagram of visual feature extraction. Figure 7 The large matrix on the left represents the input image, where each value represents the value of a pixel. The smaller matrix in the middle is the aforementioned W_1W_2 W_3, which represents the filters, and each value is a parameter of the filter. Figure 7 A portion of a medium-to-large matrix is ​​convolved with a smaller matrix, i.e. Figure 7 The formula in the formula ultimately yields the result 3, which corresponds to the value in the result matrix on the right.

[0163] Since the large matrix can be divided into 3×3 3×3 regions and convolved with the filter respectively, the result is 3×3. The n_1 mentioned earlier actually refers to the number of filter layers, that is, n_1 different intermediate small matrices are convolved with the original image respectively, resulting in a 3×3×n_1 matrix.

[0164] In this application, CNN can be used to extract visual features from short videos.

[0165] First, in order to extract visual features from short videos, it is necessary to extract keyframes from the video. Keyframes are static images in the video that represent the content at a certain moment, and they can be selected using various methods. This application uses a method of selecting one frame at fixed time intervals for extraction.

[0166] The input to a CNN consists of keyframes extracted from short videos. These keyframes need to be resized to fit the model's input requirements. The selected keyframes are then fed into the CNN's convolutional layers. The main function of the convolutional layers is to extract local features of the image, such as edges, corners, and textures. Each convolutional layer contains multiple convolutional kernels, each capturing a specific feature of the image. The convolutional kernels perform a sliding window operation on the input image, calculating the dot product between the kernel and a local region of the input image to obtain a convolutional feature map.

[0167] The output of a CNN is a high-dimensional feature vector, which is an abstract representation of the input keyframes and contains the visual information within those keyframes. This feature vector can be used for subsequent similarity calculations or quality scoring.

[0168] 1.2.2 Extracting Keyword Features Using the TF-IDF Algorithm

[0169] TF-IDF (Term Frequency-Inverse Document Frequency) is a text mining method used to measure the importance of a word in a document. The basic principle of TF-IDF is that the importance of a word in a document is proportional to its frequency of occurrence in the document and its rarity in the entire corpus.

[0170] Input: The input to the TF-IDF algorithm is typically one or more pieces of text. This text can be an article, a sentence, a dialogue, etc. When processing short videos, the input text is usually the video's metadata, such as title, description, and tags.

[0171] Output: The TF-IDF algorithm outputs the TF-IDF value for each word, forming a word vector. The TF-IDF value of each word indicates its importance in the text. The length of the word vector is equal to the size of the vocabulary, and each element represents the TF-IDF value of the corresponding word. For words not in the vocabulary, their TF-IDF value is typically set to 0.

[0172] For example, suppose there is a short video titled "Amazing underwater diving in Southeast Asia" and described as "Diving in Southeast Asia is an unforgettable experience. The underwater world is full of colorful fishes and corals." The system provided in this application first merges the title and description, and then processes the text using the TF-IDF algorithm. Assuming the vocabulary contains words such as "amazing," "underwater," "diving," and "Asia," the output word vectors might look like this:

[0173]

[0174] This word vector represents the importance of words such as "amazing," "underwater," "diving," and "Asia" in the text. The system provided in this application uses this word vector as a keyword feature of short videos for subsequent feature fusion and similarity calculation.

[0175] 1.3 Short Video Content Quality Scoring Module

[0176] The main task of the short video content quality scoring module is to score user-uploaded short videos and evaluate their quality. This module first uses a CNN algorithm to extract visual features from the user-uploaded short videos. These features capture visual information such as color, texture, and shape. Then, these visual features are input into a pre-trained neural network.

[0177] This neural network is a multi-classification model capable of classifying videos into different score ranges from 0 to 100, where 0 represents the worst video quality and 100 represents the best. A large amount of pre-collected, pre-scored short video data was used during the training of the neural network. This data includes short videos of various types and qualities, which helps the neural network learn the relationship between video quality and visual features.

[0178] The short video content quality scoring module can assign a quality score to each user's uploaded short video. This score can serve as a reference indicator of the user's understanding of the interest group's topic, and is then input into the user admission comprehensive evaluation module for final assessment.

[0179] 1.4 Short Video Content and Theme Similarity Determination Module

[0180] like Figure 8 The diagram shows the flowchart for calculating keyword similarity scores and visual feature similarity scores. The main task of the short video content and topic similarity determination module is to evaluate the similarity between user-uploaded short video content and interest group topics. This module processes video features and keyword features separately, calculates similarity scores, and normalizes the similarity scores to a range of 0-100.

[0181] First, the module compares the visual features extracted from user-uploaded videos with those extracted from administrator-uploaded videos. It calculates the similarity between these two sets of features using Euclidean distance, and then normalizes the similarity score to a range of 0-100 to obtain the visual similarity score.

[0182] Secondly, this module compares the keyword vectors extracted from the video metadata uploaded by users with the keyword vectors extracted from the videos uploaded by administrators. It also uses Euclidean distance to calculate the similarity between these two sets of keyword vectors and normalizes the similarity score to a range of 0-100, thus obtaining the keyword similarity score.

[0183] The short video content and topic similarity assessment module generates two sets of similarity scores, representing the degree of similarity between the user-uploaded short video content and the interest group's topic in terms of visual and keyword features. These similarity scores serve as reference indicators for evaluating the relevance of the user to the interest group's topic and are input into the user admission comprehensive evaluation module for final assessment.

[0184] This application calculates the Euclidean distance between the video vector and the keyword vector respectively, and then normalizes the result to the interval of 0-100 as the final similarity score. The Euclidean distance between vector A and vector B is calculated as shown in the aforementioned formula (1).

[0185] 1.5 User Access Comprehensive Evaluation Module

[0186] The user admission comprehensive assessment module is responsible for comprehensively evaluating a user's eligibility to join an interest group and determining whether the user is suitable. This module does not involve machine learning algorithms; instead, it uses code logic to perform weighted averaging and other processing on the scores obtained from the previous modules to arrive at the user's comprehensive admission score.

[0187] In addition to score-based evaluation, the user admission comprehensive assessment module also includes a risk control mechanism to prevent malicious application flooding and application fraud. Risk control is primarily based on the user's application history and interest group interaction history data. If a user is found to have frequent applications, unrelated application topics, an excessively low average admission rate, or has been previously complained about by administrators or flagged as a malicious user by the system, the system can directly reject the user's application.

[0188] After completing the risk control assessment, the user admission comprehensive assessment module will select users whose scores reach the preset threshold. The system provided in this application will then send these users' applications to the administrator for final approval. This module ensures the quality of admission to interest groups and improves the effectiveness of communication and discussion within these groups.

[0189] 1.6 Innovations and Challenges

[0190] Innovation points:

[0191] 1. Short Video-Based Admission Assessment Mechanism: Traditional admission methods for interest groups typically rely on user self-reported information or the subjective judgment of administrators. This application proposes an admission assessment mechanism based on user-uploaded short videos. By analyzing the content of user-submitted short videos, it is possible to more objectively and accurately assess users' interest in and understanding of the interest group's topic. This method helps improve the quality of admission to interest groups, reduces the review burden on administrators, and enhances the effectiveness of communication and discussion within the interest groups. Furthermore, this short video-based assessment mechanism can also enhance user participation and encourage users to actively showcase their interests and talents.

[0192] 2. The user admission comprehensive evaluation module is a key component of the system provided in this application. It is responsible for comprehensively evaluating the eligibility of users and determining their suitability for joining the interest group. This module does not involve machine learning algorithms; instead, it uses code logic to perform weighted averaging and other processing on the scores obtained from the previous modules to arrive at the user's comprehensive admission score. In addition to score evaluation, this module also has a risk control mechanism to prevent malicious application flooding and application fraud. This method ensures the quality of admission to the interest group and reduces the review burden on administrators.

[0193] difficulty:

[0194] 1. Efficient and Accurate Keyframe Extraction: To extract visual features from short videos, it is necessary to first extract the keyframes. The selection of keyframes has a significant impact on the effectiveness of visual feature extraction. How to accurately extract representative keyframes in a short time, reflecting the main content of the video while maintaining computational efficiency, is a challenge. Therefore, this application may employ different keyframe extraction strategies for different types of short videos.

[0195] 2. Training and Generalization Ability of the Neural Network Model: A pre-trained neural network model was used in the short video content quality scoring module. To obtain accurate quality scores, this classification model needs to be trained on a large amount of pre-collected and pre-scored short video data. However, the diversity and complexity of short video data may lead to insufficient generalization ability of the model. How to select appropriate network structure, loss function, and optimization algorithm, as well as how to perform effective data augmentation and model regularization to improve the model's generalization ability, is a challenging problem.

[0196] It is understood that the technical solutions provided in this application, in addition to using user-uploaded short videos to analyze users' understanding and matching degree of interest group topics, can also collect and analyze the behavior and interest preferences of each participant. The system provided in this application can understand users' interests and preferences by analyzing their behavior on the platform (such as search history, browsing history, likes, shares, comments, etc.). Furthermore, the system provided in this application can also allow users to fill out questionnaires about their interests and professional knowledge when applying to join an interest group.

[0197] The beneficial effects of the technical solution in this application include:

[0198] The interest group admission approval system proposed in this application is innovative in its mechanism, providing a new, fast, effective and accurate method for the management and admission of interest groups.

[0199] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0200] Based on the same inventive concept, this application also provides a communication group processing apparatus for implementing the communication group processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of the one or more communication group processing apparatus embodiments provided below can be found in the limitations of the communication group processing method described above, and will not be repeated here.

[0201] In one embodiment, such as Figure 9 As shown, a processing device for an exchange group is provided, including: an acquisition module 902 and a determination module 904, wherein:

[0202] The acquisition module 902 is used to acquire the visual features and keyword features corresponding to the video submitted by the target object.

[0203] The determination module 904 is used to determine the quality score of the video based on visual features; determine the visual similarity between the visual features and the preset visual features of the communication group; determine the keyword similarity between the keyword features and the preset word features of the communication group; and determine the admission information for the target object to join the communication group based on the video quality score, visual similarity, and keyword similarity.

[0204] In one embodiment, the apparatus further includes: an extraction module, configured to extract video content features of the video and use the video content features as the visual features; extract keyword features corresponding to the metadata of the video; the metadata being information used to describe the content and attributes of the video.

[0205] In one embodiment, the video content features include feature vectors; the apparatus further includes: an extraction module for extracting keyframes from the video; an adjustment module for adjusting the size of the keyframes to obtain adjusted keyframes; and an extraction module for extracting feature vectors from the adjusted keyframes and using the feature vectors as the visual features.

[0206] In one embodiment, the extraction module is further configured to extract features from the adjusted keyframes through the convolutional layer of a video feature extraction model to obtain local feature vectors of the adjusted keyframes.

[0207] In one embodiment, the meta-information includes at least the title information and descriptive statement of the video; the device further includes: a combination module, configured to combine the title information and the descriptive statement to obtain combined text information; the determination module is further configured to determine the word vector of the combined text information and use the word vector as the keyword feature corresponding to the meta-information.

[0208] In one embodiment, the apparatus further includes a processing module for classifying the visual features using a classification model to obtain a quality score for the video.

[0209] In one embodiment, the apparatus further includes: a calculation module for calculating an initial visual similarity between the visual feature and the preset visual feature; and a processing module for normalizing the initial visual similarity to obtain the visual similarity.

[0210] In one embodiment, the apparatus further includes: a calculation module, configured to calculate the cosine similarity between the keyword features and the preset word features; or to calculate the Euclidean distance between the keyword features and the preset word features; and a processing module, configured to normalize the cosine similarity or the Euclidean distance to obtain the keyword similarity.

[0211] In one embodiment, the determining module is further configured to determine a comprehensive score based on the weighting coefficient, the video quality score, the visual similarity, and the keyword similarity; and to determine the admission information for the target object to join the communication group based on the comprehensive score and risk control data.

[0212] In one embodiment, the weighting coefficients include a first weighting coefficient, a second weighting coefficient, and a third weighting coefficient; the determining module is further configured to determine, based on the characteristics of the communication group, the first weighting coefficient corresponding to the visual similarity and the second weighting coefficient corresponding to the keyword similarity; determine a video matching score based on the visual similarity, the keyword similarity, the first weighting coefficient, and the second weighting coefficient; determine a target quality score based on the video quality score and the third weighting coefficient; and determine the comprehensive score based on the video matching score and the target quality score.

[0213] In one embodiment, the determining module is further configured to determine a first product between the visual similarity and the first weight coefficient; determine a second product between the keyword similarity and the second weight coefficient; and use the sum of the first product and the second product as the video matching score.

[0214] In one embodiment, the apparatus further includes: a generation module, configured to generate admission information for rejecting the target object from joining the communication group when the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data meets the risk control conditions; and a determination module, configured to determine the admission information for the target object to join the communication group based on the comprehensive score when the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data does not meet the risk control conditions.

[0215] In one embodiment, the apparatus further includes: a comparison module, configured to compare the comprehensive score with a preset score to obtain a comparison result; a sending module, configured to send the application information of the target object corresponding to the comprehensive score to an administrator device when the comparison result indicates that the comprehensive score is greater than the preset score, so that the administrator device determines the admission information of the target object to join the communication group; a sorting module, configured to sort the comprehensive score with the comprehensive scores of other objects to obtain a sorted comprehensive score; and a determining module, further configured to determine the admission information of the target object to join the communication group from the sorted comprehensive scores.

[0216] Each module in the aforementioned communication group's processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the operations corresponding to each module.

[0217] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores communication group processing data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a communication group processing method.

[0218] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0219] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0220] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0221] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0222] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0223] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0224] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for processing chat groups, characterized in that, The method includes: Obtain the visual and keyword features corresponding to the video submitted by the target object; The quality score of the video is determined based on the visual features; Determine the visual similarity between the visual features and the preset visual features of the communication group; Determine the keyword similarity between the keyword features and the preset word features of the communication group; Based on the video quality score, the visual similarity, and the keyword similarity, the admission information for the target object to join the communication group is determined.

2. The method according to claim 1, characterized in that, The acquisition of visual features and keyword features corresponding to the video submitted by the target object includes: Extract the video content features from the video and use the video content features as the visual features; Extract keyword features corresponding to the metadata of the video; the metadata is information used to describe the content and attributes of the video.

3. The method according to claim 2, characterized in that, The video content features include feature vectors; The step of extracting video content features from the video and using the video content features as the visual features includes: Extract keyframes from the video; The size of the keyframe is adjusted to obtain the adjusted keyframe; Extract the feature vectors of the adjusted keyframes and use the feature vectors as the visual features.

4. The method according to claim 3, characterized in that, The extraction of the feature vector of the adjusted keyframe includes: The adjusted keyframes are subjected to feature extraction by the convolutional layer of the video feature extraction model to obtain the local feature vectors of the adjusted keyframes.

5. The method according to claim 2, characterized in that, The metadata includes at least the video's title and description. The extraction of keyword features corresponding to the metadata of the video includes: The title information and the description statement are combined to obtain combined text information; Determine the word vectors of the combined text information, and use the word vectors as the keyword features corresponding to the meta-information.

6. The method according to claim 1, characterized in that, Determining the quality score of the video based on the visual features includes: The visual features are classified using a classification model to obtain the quality score of the video.

7. The method according to claim 1, characterized in that, Determining the visual similarity between the visual features and preset visual features of the communication group includes: Calculate the initial visual similarity between the visual feature and the preset visual feature; The initial visual similarity is normalized to obtain the visual similarity.

8. The method according to claim 1, characterized in that, Determining the keyword similarity between the keyword features and the preset word features of the communication group includes: Calculate the cosine similarity between the keyword features and the preset word features; or Calculate the Euclidean distance between the keyword features and the preset word features; The cosine similarity or the Euclidean distance is normalized to obtain the keyword similarity.

9. The method according to claim 1, characterized in that, The process of determining the access information for the target object to join the communication group based on the video quality score, visual similarity, and keyword similarity includes: A comprehensive score is determined based on the weighting coefficients, the video quality score, the visual similarity, and the keyword similarity. Based on the comprehensive score and risk control data, the admission information for the target object to join the communication group is determined.

10. The method according to claim 9, characterized in that, The weighting coefficients include a first weighting coefficient, a second weighting coefficient, and a third weighting coefficient; the method further includes: Based on the characteristics of the communication group, a first weight coefficient corresponding to the visual similarity and a second weight coefficient corresponding to the keyword similarity are determined. The determination of the comprehensive score based on weighting coefficients, video quality scores, visual similarity, and keyword similarity includes: The video matching score is determined based on the visual similarity, the keyword similarity, the first weight coefficient, and the second weight coefficient. The target quality score is determined based on the video quality score and the third weighting coefficient. The comprehensive score is determined based on the video matching score and the target quality score.

11. The method according to claim 10, characterized in that, The process of determining the video matching score based on the visual similarity, the keyword similarity, the first weight coefficient, and the second weight coefficient includes: Determine the first product between the visual similarity and the first weighting coefficient; Determine the second product between the keyword similarity and the second weight coefficient; The sum of the first product and the second product is used as the video matching score.

12. The method according to claim 9, characterized in that, The process of determining the access information for the target object to join the communication group based on the comprehensive score and risk control data includes: When the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data meets the risk control conditions, access information is generated to deny the target object from joining the communication group; When the correlation between the historical application success rate, historical application frequency, and historical application topic in the risk control data does not meet the risk control conditions, the admission information for the target object to join the communication group is determined based on the comprehensive score.

13. The method according to claim 12, characterized in that, The process of determining the admission information for the target object to join the communication group based on the comprehensive score includes: The comprehensive score is compared with a preset score to obtain a comparison result; when the comparison result indicates that the comprehensive score is greater than the preset score, the application information of the target object corresponding to the comprehensive score is sent to the administrator device, so that the administrator device can determine the access information for the target object to join the communication group; or The overall score is sorted with the overall scores of other objects to obtain a sorted overall score; the admission information for the target object to join the communication group is determined from the sorted overall scores.

14. A processing device for a communication group, characterized in that, The device includes: The acquisition module is used to acquire the visual features and keyword features corresponding to the video submitted by the target object; The determination module is used to determine the quality score of the video based on the visual features; determine the visual similarity between the visual features and the preset visual features of the communication group; determine the keyword similarity between the keyword features and the preset word features of the communication group; and determine the admission information for the target object to join the communication group based on the video quality score, the visual similarity, and the keyword similarity.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.