Association degree determination method and program product

By using ASR and OCR algorithms to identify video text features, and combining the relevance recognition model of store features, the problem of low video quality and correlation is solved, and the exposure rate and product display effect of online stores are improved.

CN119991207APending Publication Date: 2025-05-13BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059007.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the video uploaded by users is of poor quality and has a low correlation with the target store, resulting in poor exposure and product display effects of online stores.

Method used

Through an associative determination method, the target video is recognized using automatic speech recognition algorithm (ASR) and optical character recognition algorithm (OCR), the text features of the video are obtained, and combined with the text features of the target store, the training correlation recognition model is used to process these features to determine the correlation between the target video and the target store.

Benefits of technology

The efficiency of determining the correlation between the target video and the target store is improved, ensuring that the video and the store’s display content are more relevant and quality, thereby improving the exposure and product display effect of online stores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991207A_ABST
    Figure CN119991207A_ABST
Patent Text Reader

Abstract

The invention provides a correlation degree determination method and a program product, and relates to the technical field of computers, and the method comprises the steps: responding to the selection of a target video by a user, displaying a target shop corresponding to the target video, responding to the selection operation of the target shop and the target video by the user, displaying a correlation degree identification result of the target shop and the target video, and the efficiency of determining the association degree between the target video and the target shop is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of the times, online consumption has gradually become the mainstream of daily consumption. Users open online stores to display and sell goods. In order to enhance the exposure of online stores and improve the display effect of products in online stores. Users often upload videos related to the target store, but the videos currently uploaded by users have problems such as poor video quality and low relevance to the target video.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0004] The present disclosure provides a method for determining a relevance and a program product, which at least determine a relevance between a target video and a target store.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a method for determining a degree of association is provided, comprising:

[0007] In response to the user's selection of a target video, a target store corresponding to the target video is displayed, and the target video is used to display the content included in the target store;

[0008] In response to the user's selection operation on the target store and the target video, the association recognition result between the target store and the target video is displayed.

[0009] In one embodiment of the present disclosure, in response to a user's selection operation of a target store and a target video, a result of identifying the degree of association between the target store and the target video is displayed, including:

[0010] Obtaining a first text feature corresponding to a target store;

[0011] Recognize the target video based on the Automatic Speech Recognition (ASR) algorithm and the Optical Character Recognition (OCR) algorithm to obtain the second text feature;

[0012] The first text feature, the second text feature, and the target video are processed based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

[0013] In one embodiment of the present disclosure, the method further includes:

[0014] The intelligent model is trained based on the constructed training samples to obtain training results;

[0015] Determine the loss function values ​​corresponding to the ITC loss function, the ITM loss function, and the LM loss function respectively based on the training results;

[0016] Accumulate the loss function values ​​corresponding to the ITC loss function, ITM loss function, and LM loss function to obtain the target loss function value;

[0017] When the target loss function value reaches a preset threshold, the training is determined to be completed and the association recognition model is obtained.

[0018] In one embodiment of the present disclosure, the method further includes:

[0019] Input the training video into the big model to obtain the training text;

[0020] The training video and the training text are determined as training samples.

[0021] In one embodiment of the present disclosure, the first text feature, the second text feature and the target video are processed based on the trained relevance recognition model to obtain a relevance recognition result between the target store and the target video, including:

[0022] Encoding the first text feature and the second text feature based on a Chinese encoder to obtain encoded target text features;

[0023] The image obtained by extracting frames from the target video is processed based on the image block embedding layer and the position embedding layer to obtain encoded image features;

[0024] Input the encoded target text features and the encoded image features into the cross attention layer to obtain the aligned target text features and image features;

[0025] The aligned target text features and image features are input into the output result generation layer to obtain the correlation recognition result between the target store and the target video.

[0026] In one embodiment of the present disclosure, the method further includes:

[0027] Cascading the first text feature and the second text feature to obtain a target text feature;

[0028] The first text feature and the second text feature are encoded based on the Chinese encoder to obtain the encoded target text feature, including:

[0029] The target text features are encoded based on the Chinese encoder to obtain the encoded target text features.

[0030] In one embodiment of the present disclosure, the target text features are encoded based on a Chinese encoder to obtain the encoded target text features, including:

[0031] Divide the target text features into blocks to obtain multiple block text features;

[0032] Add classification marks to block text features to record the global features of block text features;

[0033] The block text features with added classification marks are encoded to obtain the encoded target text features.

[0034] In one embodiment of the present disclosure, the method further includes:

[0035] Extract frames from the target video to obtain an extracted frame image;

[0036] Obtain image parameters of multiple frame-sampling images;

[0037] The image parameters are averaged to obtain the averaged image parameters;

[0038] The extracted frame images are adjusted based on the averaged image parameters to obtain an image.

[0039] In one embodiment of the present disclosure, an image obtained by extracting frames from a target video is processed based on an image block embedding layer and a position embedding layer to obtain encoded image features, including:

[0040] The image is processed into blocks to obtain multiple image blocks;

[0041] Add classification tags to image blocks to record global features of the image blocks;

[0042] The image blocks with added classification marks are encoded to obtain the encoded image features.

[0043] According to another aspect of the present disclosure, there is provided a device for determining a degree of association, comprising:

[0044] A first display module, configured to display a target store corresponding to the target video in response to a user's selection of the target video, wherein the target video is used to display the content included in the target store;

[0045] The second display module is used to display the association recognition result between the target store and the target video in response to the user's selection operation on the target store and the target video.

[0046] In one embodiment of the present disclosure, the second display module includes:

[0047] An acquisition unit, used to acquire a first text feature corresponding to a target store;

[0048] A recognition unit, used to recognize the target video based on an automatic speech recognition algorithm ASR and an optical character recognition algorithm OCR to obtain a second text feature;

[0049] The processing unit is used to process the first text feature, the second text feature and the target video based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

[0050] In one embodiment of the present disclosure, the device further includes:

[0051] The training module is used to train the intelligent model based on the constructed training samples to obtain the training results;

[0052] A first determination module is used to determine the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function respectively based on the training results;

[0053] An accumulation module is used to accumulate the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function to obtain the target loss function value;

[0054] The second determination module is used to determine the end of training and obtain the association recognition model when the target loss function value reaches a preset threshold.

[0055] In one embodiment of the present disclosure, the device further includes:

[0056] The input module is used to input the training video into the large model to obtain the training text;

[0057] The third determination module is used to determine the training video and the training text as training samples.

[0058] In one embodiment of the present disclosure, a processing unit includes:

[0059] An encoding subunit, used for encoding the first text feature and the second text feature based on a Chinese encoder to obtain an encoded target text feature;

[0060] A processing subunit, used for processing an image obtained by extracting frames from a target video based on an image block embedding layer and a position embedding layer to obtain encoded image features;

[0061] A first input subunit is used to input the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features;

[0062] The second input subunit is used to input the aligned target text features and image features into the output result generation layer to obtain the association recognition result between the target store and the target video.

[0063] In one embodiment of the present disclosure, the device further includes:

[0064] A cascading module, used for cascading the first text feature and the second text feature to obtain a target text feature;

[0065] Coding subunit, including:

[0066] The first encoding component is used to encode the target text feature based on the Chinese encoder to obtain the encoded target text feature.

[0067] In one embodiment of the present disclosure, the encoding subunit includes:

[0068] A block component is used to block the target text feature to obtain multiple block text features;

[0069] The first recording element is used to add classification marks to the block text features to record the global features of the block text features;

[0070] The second encoding element is used to encode the block text features with the classification mark added to obtain the encoded target text features.

[0071] In one embodiment of the present disclosure, the device further includes:

[0072] A frame extraction module is used to extract frames of a target video to obtain a frame extraction image;

[0073] An acquisition module, used for acquiring image parameters of a plurality of frame-sampling images;

[0074] A calculation module, used for averaging the image parameters to obtain averaged image parameters;

[0075] The adjustment module is used to adjust the extracted frame image based on the averaged image parameters to obtain an image.

[0076] In one embodiment of the present disclosure, the processing subunit includes:

[0077] A processing element, used for processing the image into blocks to obtain a plurality of image blocks;

[0078] A second recording element is used to add a classification mark to the image block to record the global features of the image block;

[0079] The third encoding element is used to encode the image block with the classification mark added to obtain the encoded image features.

[0080] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above-mentioned methods for determining a degree of association by executing the executable instructions.

[0081] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned methods for determining a degree of association is implemented.

[0082] According to another aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program or computer instructions. The computer program or computer instructions are loaded and executed by a processor to enable a computer to implement any of the above-mentioned methods for determining a degree of association.

[0083] In the disclosed embodiment, in response to a user's selection of a target video, a target store corresponding to the target video is displayed, and in response to a user's selection operation of the target store and the target video, an identification result of the association between the target store and the target video is displayed, thereby improving the efficiency of determining the association between the target video and the target store.

[0084] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0086] Figure 1 A schematic diagram of the structure of a correlation determination system in an embodiment of the present disclosure is shown.

[0087] Figure 2 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0088] Figure 3 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0089] Figure 4 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0090] Figure 5 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0091] Figure 6A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0092] Figure 7 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0093] Figure 8 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown.

[0094] Fig. 9 A schematic diagram of a device for determining a degree of association in an embodiment of the present disclosure is shown.

[0095] Fig.10 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown.

[0096] Fig.11 A schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0097] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0098] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0099] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0100] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0101] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0102] In some application scenarios, the target video may be a product display video, a promotional activity video, a user evaluation video, a product usage demonstration video, and a store display video of a target store. The present disclosure does not specifically limit the application scenario. The present disclosure may be applied to any scenario including a store and a video.

[0103] It should be pointed out that, in the absence of conflict, the embodiments of the present disclosure and the technical features therein may be combined with each other.

[0104] Figure 1 A schematic diagram of the structure of a relevance determination system in an embodiment of the present disclosure is shown, and the system can be applied to the relevance determination method and relevance determination device in various embodiments of the present disclosure.

[0105] like Figure 1 As shown, the association degree determination system 10 may include a front-end display module 101 and a back-end processing module 102 .

[0106] In some embodiments, the front-end display module 101 and the back-end processing module 102 may be located on different devices, and the front-end display module 101 may be a module on an electronic device with data acquisition capability. The back-end processing module 102 may be a module on an electronic device with data processing capability, such as a computer, a server, etc. The front-end display module 101 and the back-end processing module 102 may be located on the same device, such as the front-end display module 101 and the back-end processing module 102 may be an input module and a processing module on a computer or a mobile phone.

[0107] The front-end display module 101 and the back-end processing module 102 are connected to each other via a network, which may be a wired network or a wireless network.

[0108] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0109] The following describes a situation where the front-end display module 101 and the back-end processing module 102 are located on two different devices.

[0110] The front-end display module 101 may be located on a terminal device, which may be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, wearable devices, augmented reality devices, virtual reality devices, etc.

[0111] Optionally, the client of the application installed in different terminal devices is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile client, a PC client, etc.

[0112] The backend processing module 102 may be located on a server, which may be a server that provides various services, such as a backend management server that provides support for devices operated by users using terminal devices. The backend management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.

[0113] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0114] Those skilled in the art will know that Figure 1 The number of the front-end display modules 101 and the back-end processing modules 102 is only illustrative, and any number of the front-end display modules 101 and the back-end processing modules 102 may be provided according to actual needs. The embodiments of the present disclosure are not limited to this.

[0115] Figure 2 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0116] S210, in response to the user's selection of a target video, displaying a target store corresponding to the target video, where the target video is used to display the content included in the target store.

[0117] In some embodiments, the target video may be a video uploaded by a user, and the user may be a user corresponding to the target store.

[0118] In some embodiments, the user who uploads the target video and the user who selects the target video may be the same user or different users.

[0119] In some embodiments, before S210, an association relationship between the target video and the target store may be established. For example, if the user who uploaded the target video and the user corresponding to the store are the same user, an association relationship between the target video and the target store may be established based on the user's identifier.

[0120] In some embodiments, after the user selects the target video, the terminal device obtains information about the user selecting the target video, generates a selection instruction based on the information about the user selecting the target video, and sends the selection instruction to the server, where the selection instruction includes the identifier of the target video. After receiving the selection instruction, the server determines the information about the target store based on the identifier of the target video and the predetermined association relationship between the target video and the target store, and sends the information about the target store to the terminal device, so that the terminal device displays the target store based on the information about the target store.

[0121] In some embodiments, in response to the user's selection of a target store, a target video corresponding to the target store may be displayed, where the target video is used to display the content contained in the target store.

[0122] S220 , in response to the user's selection operation on the target store and the target video, displaying the association recognition result between the target store and the target video.

[0123] In some embodiments, the association recognition result may be expressed as similarity, and the larger the value corresponding to the similarity, the higher the association.

[0124] In some embodiments, the association recognition result between the target store and the target video can be used to indicate the matching degree between the target store and the target video. For example, if the store name and store category name of the target store match the video content in the target video, the association recognition result indicates a high association, and if the store name and store category name of the target store do not match the video content in the target video, the association recognition result indicates a low association.

[0125] In some embodiments, after the user selects the target video and the target store, the terminal device obtains the information of the user selecting the target video and the target store, generates a selection signaling based on the information of the user selecting the target video and the target store, and sends the selection signaling to the server, the selection signaling including the information of the target store and the target video. After receiving the selection signaling, the server processes the relevant information of the target store and the target video to obtain the correlation recognition result of the target store and the target video.

[0126] In the disclosed embodiment, in response to a user's selection of a target video, a target store corresponding to the target video is displayed, and in response to a user's selection operation of the target store and the target video, an identification result of the association between the target store and the target video is displayed, thereby improving the efficiency of determining the association between the target video and the target store.

[0127] Figure 3 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 3 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0128] S310, in response to the user's selection of a target video, displaying a target store corresponding to the target video, where the target video is used to display the content included in the target store.

[0129] S320: Obtain a first text feature corresponding to the target store.

[0130] In some embodiments, the first text feature corresponding to the target store may include all text features related to the target store and text features converted from all other features related to the target store. Exemplarily, the first text feature may include the name of the target store, the name of the category to which the target store belongs, the name of the products included in the target store, and text data converted from other audio information.

[0131] In some embodiments, the first text feature of the target store may be acquired based on a variety of methods such as a crawler algorithm and user input.

[0132] S330, recognizing the target video based on an automatic speech recognition algorithm ASR and an optical character recognition algorithm OCR to obtain a second text feature.

[0133] In some embodiments, the automatic speech recognition algorithm ASR is a technology that converts human speech signals into text. A target video can be input to obtain text information corresponding to the target video.

[0134] In some embodiments, an optical character recognition (OCR) algorithm is a technology that converts text in an image into computer-editable text. The text information in the image is captured by an electronic device (such as a scanner or a digital camera), and then image processing technology is used to detect dark and light patterns in the image to determine the shape of the text. Then, these shapes are translated into a text format that can be understood by the computer through a character recognition algorithm.

[0135] In some embodiments, the target video is input into ASR to obtain first text information, and the target video is input into OCR to obtain second text information. The second text feature may include text type data obtained by concatenating the first text information and the second text information.

[0136] In some embodiments, the cascading of the first text information and the second text information includes merging the data in any manner while retaining the first text information and the second text information to obtain the second text feature.

[0137] S340, processing the first text feature, the second text feature and the target video based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

[0138] In the disclosed method, a first text feature of a target store is obtained, a second text feature corresponding to a target video is obtained based on different algorithms, the first text feature, the second text feature and the target video are processed based on a trained association recognition model to obtain an association recognition result of the target video. Acquiring text features of the target video through a variety of algorithms can ensure that the acquired data is relatively comprehensive, and further improve the accuracy of the recognition results.

[0139] Figure 4 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 4 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0140] S410: Train the intelligent model based on the constructed training samples to obtain training results.

[0141] In some embodiments, the training samples may include training videos and training texts.

[0142] In some embodiments, before S410, the method may further include:

[0143] Input the training video into the big model to obtain the training text;

[0144] The training video and the training text are determined as training samples.

[0145] In some embodiments, the large model can be a machine learning model with large-scale parameters and complex computational structures, which are usually constructed by deep neural networks. Exemplarily, the large model can include a QWEN model.

[0146] Exemplarily, the training video may be frame-extracted to obtain multiple images corresponding to the training video, and then the images may be input into the large model and the instruction "describe the input image" may be input. After the instruction is input, the training text corresponding to the training video may be obtained.

[0147] In some embodiments, determining training samples based on a large model can efficiently and quickly obtain training texts, thereby improving the efficiency of determining relevance.

[0148] S420, determining loss function values ​​corresponding to the ITC loss function, the ITM loss function, and the LM loss function, respectively, based on the training results.

[0149] In some embodiments, ITC loss, Image-Text Contrastive Loss, is a classic loss function in multimodality, which allows matched image-text pairs to have a higher similarity expression. ITM loss, Image-Text Matching Loss, learns a binary classification loss function for fine-grained matching of image-text pairs. LM loss, Language Modeling Loss, learns how to generate a coherent text description from a given image, and uses a cross entropy cost function to maximize the corresponding text probability in an autoregressive manner.

[0150] S430, accumulating the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function respectively to obtain a target loss function value.

[0151] In some embodiments, accumulating multiple loss function values ​​may include adding the multiple loss function values ​​to obtain a target loss function value.

[0152] In some embodiments, the loss function values ​​corresponding to multiple different loss functions are superimposed to determine whether the intelligent model has been trained, which can better align the similarity between the image and the Chinese features, so that the trained intelligent model can determine the recognition results more accurately when determining the association recognition results.

[0153] S440, when the target loss function value reaches a preset threshold, determine that the training is completed and obtain the association recognition model.

[0154] In some embodiments, the preset threshold may be determined by a user, which is not specifically limited in the present disclosure.

[0155] S450, processing the first text feature, the second text feature and the target video based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

[0156] In the disclosed embodiment, the intelligent model is trained through training samples, and the target loss function value is determined based on multiple loss functions. When the target loss function value reaches a preset threshold, a trained correlation recognition model is obtained. The trained correlation recognition model meets the requirements of multiple aspects and improves the accuracy of the correlation recognition model.

[0157] Figure 5 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 5 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0158] S510: Obtain a first text feature corresponding to the target store.

[0159] In some embodiments, the first text feature includes a Chinese description related to the target store.

[0160] S520, recognizing the target video based on an automatic speech recognition algorithm ASR and an optical character recognition algorithm OCR to obtain a second text feature.

[0161] In some embodiments, the second text feature may include a text feature obtained by converting audio and images in the target video.

[0162] In some embodiments, identifying the target video based on the ASR algorithm and OCR respectively can enable the obtained second text feature to include more comprehensive information in the target video.

[0163] S530, encoding the first text feature and the second text feature based on a Chinese encoder to obtain encoded target text features.

[0164] In some embodiments, the Chinese encoder may convert Chinese text or characters into a format that can be understood and processed by the intelligent model.

[0165] In some embodiments, the Chinese encoder in the present disclosure may include an embedding layer, a convolutional layer, a recurrent layer, and an attention mechanism.

[0166] For example, the embedding layer can convert Chinese characters or words into vector representations, which can capture the semantic relationship between characters or words. The convolution layer is used to extract local features in the text, such as n-gram features. The recurrent layer is used to capture the dependencies between words in a sentence, and the recurrent layer can include RNN, LSTM, and GRU. The attention mechanism is used to calculate the correlation between different parts of the input data and dynamically adjust the weight of each part. This helps the model better focus on the important information in the input data.

[0167] In some embodiments, encoding the first text feature and the second text feature based on the Chinese encoder avoids the problem of information loss caused by language conversion when using other encoders. On the other hand, encoding the first text feature and the second text feature makes the acquired text features more comprehensive, thereby improving the accuracy of the association recognition results.

[0168] In some embodiments, before S530, the method may further include: performing a cascade operation on the first text feature and the second text feature. The method of the cascade operation may be the same as the method of the cascade operation in the above embodiment, which will not be described in detail herein.

[0169] In some embodiments, the target text feature may include a text feature represented in a vector form.

[0170] S540, processing the image obtained by extracting frames from the target video based on the image block embedding layer and the position embedding layer to obtain encoded image features.

[0171] In some embodiments, S540 may include: dividing the image obtained by frame extraction into blocks, and encoding the divided image to record the global features of the image.

[0172] In some embodiments, the images obtained by extracting frames from the target video are divided into blocks and encoded, so that text features and images can be aligned, thereby improving the accuracy of the associated recognition results.

[0173] S550, inputting the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features.

[0174] In some embodiments, the cross attention layer is a special attention mechanism that allows the model to focus on relevant information in another input sequence while processing one input sequence. The core of this mechanism is to calculate the similarity between two different input sequences and assign weights to each element based on the similarity, thereby achieving effective fusion of information.

[0175] In some embodiments, the encoded target text features and the encoded image features may be input into the self-attention layer before the encoded target text features and the encoded image features are input into the cross-self-attention layer.

[0176] In some embodiments, the commonly added cross-attention layer can align target text features and image features, improving the expressiveness of the model and the accuracy of task processing.

[0177] S560, inputting the aligned target text features and image features into the output result generation layer to obtain the correlation recognition result between the target store and the target video.

[0178] In the disclosed embodiment, a first text feature is determined, the first text feature and the second text feature are encoded based on a Chinese encoder to obtain an encoded target text feature, a target video is framed based on an image block embedding layer and a position embedding layer to obtain an encoded image feature, the target text feature and the image feature are aligned based on a cross-attention layer, the aligned target text feature and image feature are input into an output result generation layer to obtain a correlation recognition result between a target store and a target video, thereby improving the accuracy of the correlation recognition result as a whole.

[0179] Figure 6 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 6 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0180] S610, dividing the target text feature into blocks to obtain a plurality of block text features.

[0181] In some embodiments, the Chinese BERT model can be used to segment the above target text features.

[0182] S620: Add classification marks to the block text features to record the global features of the block text features.

[0183] In some embodiments, a [cls] token may be added as a classification tag.

[0184] S630, encoding the block text features with the added classification mark to obtain encoded target text features.

[0185] S640, processing an image obtained by extracting frames from the target video based on the image block embedding layer and the position embedding layer to obtain encoded image features;

[0186] S650, inputting the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features;

[0187] S660, inputting the aligned target text features and image features into the output result generation layer to obtain the correlation recognition result between the target store and the target video.

[0188] In the disclosed embodiment, by dividing and marking the target text features, the target text features can be aligned with the image features while retaining the global features of the target features, thereby improving the accuracy of the association recognition result.

[0189] Figure 7 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 7 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0190] S710, encoding the first text feature and the second text feature based on a Chinese encoder to obtain encoded target text features;

[0191] S720, extract frames from the target video to obtain an extracted frame image;

[0192] S730, obtaining image parameters of a plurality of frame-sampling images.

[0193] In some embodiments, image parameters may include: contrast, saturation, brightness, sharpness, hue, and hue.

[0194] In some embodiments, the image parameters of the image may be acquired based on image processing software, programming languages, command line tools, and the like.

[0195] S740, performing performance averaging on the image parameters to obtain averaged image parameters.

[0196] In some embodiments, averaging the image parameters may include summing and averaging the image parameters based on a mathematical operation.

[0197] Exemplarily, the image parameters of each category may be summed and then divided by the number of images to obtain the image parameters.

[0198] S750, adjusting the frame-sampling image based on the averaged image parameters to obtain an image.

[0199] In some embodiments, adjusting the extracted image based on the averaged image parameters may include setting all image parameters of the extracted image to the averaged image parameters.

[0200] S760, processing an image obtained by extracting frames from the target video based on the image block embedding layer and the position embedding layer to obtain encoded image features;

[0201] S770, inputting the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features;

[0202] S780, inputting the aligned target text features and image features into the output result generation layer to obtain the correlation recognition result between the target store and the target video.

[0203] In the disclosed embodiment, the frame-sampling image is adjusted based on the averaged image parameters, and the processed image is then subjected to subsequent processing, so that the image parameters of different frame-sampling images are the same, thereby reducing the difficulty of association recognition and improving the efficiency of association recognition.

[0204] Figure 8 A flow chart of a method for determining a degree of association in an embodiment of the present disclosure is shown as follows: Figure 8 As shown, the method for determining the degree of association in the embodiment of the present disclosure may include:

[0205] S810, encoding the first text feature and the second text feature based on a Chinese encoder to obtain encoded target text features;

[0206] S820, dividing the image into blocks to obtain a plurality of image blocks.

[0207] In some embodiments, the image can be processed in blocks based on the Vision Transformer (ViT) architecture, the Transformer encoder module, and the MLP classification module.

[0208] S830: Add a classification tag to the image block to record the global features of the image block.

[0209] In some embodiments, a [cls] token may be added to record global features, wherein the [cls] token may be used as a position code to retain the spatial features of the image blocks.

[0210] S840, encoding the image block with the classification mark added thereto to obtain encoded image features;

[0211] S850, inputting the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features;

[0212] S860, inputting the aligned target text features and image features into the output result generation layer to obtain the correlation recognition result between the target store and the target video.

[0213] In the disclosed embodiment, the operations of dividing the image into blocks, adding classification marks, and encoding can align the image of the divided blocks with the text features of the divided blocks, thereby reducing the difficulty of association recognition.

[0214] In order to explain the present disclosure in detail, the present disclosure also provides an exemplary embodiment. The following is a description of the exemplary embodiment.

[0215] Obtain a video, extract frames from the video based on a video extraction tool, obtain video extraction images, and input the obtained multiple video extraction images into a large model, such as QWEN-VL. Input the instruction "describe the above video extraction images" to the large model. After obtaining the text information output by the large model, use the above text information and video extraction images as training samples to train the association recognition model. When the target loss function value of the model reaches a preset threshold, determine that the training is completed and obtain the trained association recognition model, wherein the target loss function value is accumulated by the loss function values ​​corresponding to the ITC loss function, the ITM loss function, and the LM loss function. After the training of the association recognition model is completed, the user can select the target video and the target store, wherein the target video is used to display the content contained in the target store. After the user selects the target video and the target store, the first text feature of the target store can be obtained by the terminal device, and the first text feature can include text information related to the target store that can be directly obtained. Then the ASR algorithm and the OCR algorithm recognize the target video to obtain the second text feature associated with the target video. The first text feature and the second text feature are cascaded to obtain the target text feature. The target text feature is encoded based on the Chinese encoder to obtain the encoded target text feature. The target video is framed based on the relevant tool to obtain the image. The above image is processed based on the image block embedding layer and the position embedding layer in the association recognition model to obtain the encoded image feature. The encoded target text feature and the encoded image feature are input into the cross attention layer to obtain the aligned target text feature and image feature. The aligned target text feature and image feature are input into the result generation layer to obtain the association recognition result. The association recognition result is displayed based on the terminal device.

[0216] Based on the same inventive concept, the present disclosure also provides a device for determining a degree of association, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.

[0217] Fig. 9 A schematic diagram of a device for determining a degree of association in an embodiment of the present disclosure is shown. Fig. 9 As shown, the association degree determination device 900 includes:

[0218] The first display module 910 is used to display a target store corresponding to the target video in response to the user's selection of the target video, where the target video is used to display the content included in the target store;

[0219] The second display module 920 is used to display the association recognition result between the target store and the target video in response to the user's selection operation on the target store and the target video.

[0220] In some embodiments, the second display module 920 includes:

[0221] An acquisition unit, used to acquire a first text feature corresponding to a target store;

[0222] A recognition unit, used to recognize the target video based on an automatic speech recognition algorithm ASR and an optical character recognition algorithm OCR to obtain a second text feature;

[0223] The processing unit is used to process the first text feature, the second text feature and the target video based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

[0224] In some embodiments, the apparatus further comprises:

[0225] The training module is used to train the intelligent model based on the constructed training samples to obtain the training results;

[0226] A first determination module is used to determine the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function respectively based on the training results;

[0227] An accumulation module is used to accumulate the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function to obtain the target loss function value;

[0228] The second determination module is used to determine the end of training and obtain the association recognition model when the target loss function value reaches a preset threshold.

[0229] In some embodiments, the apparatus further comprises:

[0230] The input module is used to input the training video into the large model to obtain the training text;

[0231] The third determination module is used to determine the training video and the training text as training samples.

[0232] In some embodiments, the processing unit includes:

[0233] An encoding subunit, used for encoding the first text feature and the second text feature based on a Chinese encoder to obtain an encoded target text feature;

[0234] A processing subunit, used for processing an image obtained by extracting frames from a target video based on an image block embedding layer and a position embedding layer to obtain encoded image features;

[0235] A first input subunit is used to input the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features;

[0236] The second input subunit is used to input the aligned target text features and image features into the output result generation layer to obtain the association recognition result between the target store and the target video.

[0237] In some embodiments, the apparatus further comprises:

[0238] A cascading module, used for cascading the first text feature and the second text feature to obtain a target text feature;

[0239] Coding subunit, including:

[0240] The first encoding component is used to encode the target text feature based on the Chinese encoder to obtain the encoded target text feature.

[0241] In some embodiments, the encoding subunit comprises:

[0242] A block component is used to block the target text feature to obtain multiple block text features;

[0243] The first recording element is used to add classification marks to the block text features to record the global features of the block text features;

[0244] The second encoding element is used to encode the block text features with the classification mark added to obtain the encoded target text features.

[0245] In some embodiments, the apparatus further comprises:

[0246] A frame extraction module is used to extract frames of a target video to obtain a frame extraction image;

[0247] An acquisition module, used for acquiring image parameters of a plurality of frame-sampling images;

[0248] A calculation module, used for averaging the image parameters to obtain averaged image parameters;

[0249] The adjustment module is used to adjust the extracted frame image based on the averaged image parameters to obtain an image.

[0250] In some embodiments, the processing subunit includes:

[0251] A processing element, used for processing the image into blocks to obtain a plurality of image blocks;

[0252] A second recording element is used to add a classification mark to the image block to record the global features of the image block;

[0253] The third encoding element is used to encode the image block with the classification mark added to obtain the encoded image features.

[0254] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".

[0255] The present disclosure provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned display methods by executing the executable instructions.

[0256] For example, refer to Fig.10 The electronic device 1000 according to this embodiment of the present disclosure is described. Fig.10 The electronic device 1000 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0257] like Fig.10 As shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 may include but are not limited to: at least one processing unit 1010, at least one storage unit 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).

[0258] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 performs the steps of various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 1010 can perform the following steps of the above method embodiment:

[0259] In response to the user's selection of a target video, a target store corresponding to the target video is displayed, and the target video is used to display the content included in the target store;

[0260] In response to the user's selection operation on the target store and the target video, the association recognition result between the target store and the target video is displayed.

[0261] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202 , and may further include a read-only storage unit (ROM) 10203 .

[0262] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0263] Bus 1030 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0264] The electronic device 1000 may also communicate with one or more external devices 1040 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1060. As shown, the network adapter 1060 communicates with other modules of the electronic device 1000 via a bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0265] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0266] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium. Fig.11 A schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure is shown. Fig.11As shown, the computer-readable storage medium 1100 stores a program product capable of implementing the above method of the present disclosure.

[0267] In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present disclosure described in the above "Specific Implementation" section of this specification.

[0268] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0269] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0270] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0271] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0272] The embodiments of the present disclosure provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the display method provided in various optional ways in any embodiment of the present disclosure.

[0273] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0274] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0275] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope of the present disclosure is indicated by the appended claims.

Claims

1. A method for determining a degree of association, characterized in that: The method comprises: In response to a user's selection of a target video, a target store corresponding to the target video is displayed, wherein the target video is used to display the content included in the target store; In response to the user's selection operation on the target store and the target video, the association recognition result between the target store and the target video is displayed.

2. The method according to claim 1, characterized in that The step of displaying the result of the association degree identification between the target store and the target video in response to the user's selection operation on the target store and the target video includes: Obtaining a first text feature corresponding to a target store; Recognize the target video based on the automatic speech recognition algorithm ASR and the optical character recognition algorithm OCR to obtain the second text feature; The first text feature, the second text feature, and the target video are processed based on the trained association recognition model to obtain an association recognition result between the target store and the target video.

3. The method according to claim 2, characterized in that The method further comprises: The intelligent model is trained based on the constructed training samples to obtain training results; Determine the loss function values ​​corresponding to the ITC loss function, the ITM loss function and the LM loss function respectively based on the training results; Accumulate the loss function values ​​corresponding to the ITC loss function, the ITM loss function, and the LM loss function to obtain a target loss function value; When the target loss function value reaches a preset threshold, the training is determined to be completed and the association recognition model is obtained.

4. The method according to claim 2, characterized in that: The method further comprises: Input the training video into the big model to obtain the training text; The training video and the training text are determined as training samples.

5. The method according to claim 2, characterized in that: The trained association recognition model processes the first text feature, the second text feature, and the target video to obtain an association recognition result between the target store and the target video, including: Encoding the first text feature and the second text feature based on a Chinese encoder to obtain encoded target text features; Processing the image obtained by extracting frames from the target video based on the image block embedding layer and the position embedding layer to obtain encoded image features; Inputting the encoded target text features and the encoded image features into a cross attention layer to obtain aligned target text features and image features; The aligned target text features and image features are input into the output result generation layer to obtain the association recognition result between the target store and the target video.

6. The method according to claim 5, characterized in that The method further comprises: Cascading the first text feature and the second text feature to obtain a target text feature; The step of encoding the first text feature and the second text feature based on the Chinese encoder to obtain the encoded target text feature includes: The target text features are encoded based on a Chinese encoder to obtain encoded target text features.

7. The method according to claim 5, characterized in that The step of encoding the target text features based on the Chinese encoder to obtain the encoded target text features includes: Divide the target text features into blocks to obtain multiple block text features; Adding classification marks to the block text features to record the global features of the block text features; The block text features with the added classification mark are encoded to obtain the encoded target text features.

8. The method according to claim 5, characterized in that The method further comprises: Extract frames from the target video to obtain an extracted frame image; Acquire image parameters of a plurality of the frame-sampling images; Performing performance averaging on the image parameters to obtain averaged image parameters; The frame-sampling image is adjusted based on the averaged image parameters to obtain the image.

9. The method according to claim 5, characterized in that The image obtained by extracting frames from the target video based on the image block embedding layer and the position embedding layer to obtain encoded image features includes: Divide the image into blocks to obtain multiple image blocks; Adding a classification mark to the image block to record the global features of the image block; The image block with the added classification mark is encoded to obtain the encoded image feature.

10. A computer program product, comprising a computer program or computer instructions, characterized in that: The computer program or the computer instruction is loaded and executed by a processor, so that the computer implements the method for determining the degree of association as described in any one of claims 1 to 9.