Data detection method and apparatus, object recommendation method and apparatus, and device and medium

By generating and combining the position information of the target sample matrix and the detection data of the eigenvalues, the learning difficulty of sparse matrices in machine learning is solved, and the accuracy of the difference in feature data and the improvement of model performance is achieved.

WO2025123839A1PCT designated stage expired Publication Date: 2025-06-19MASHANG CONSUMER FINANCE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120137
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-09-20
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In machine learning, training data based on sparse matrix increases the learning difficulty of the model and reduces the performance of the model, and it is difficult for the prior art to accurately identify differences between feature data.

Method used

By acquiring the target sample matrix to be detected, the first detection data and the second detection data are generated based on the position information and characteristic values ​​of the sample characteristics, and whether the two detection matrices are sparse matrices, thereby identifying the differences between sample characteristics.

Benefits of technology

Accurate detection of the sparsity of the target sample matrix is ​​achieved, and the differences between sample features are identified, thereby reducing the difficulty of model learning and improving model performance in machine learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120137_19062025_PF_FP_ABST
    Figure CN2024120137_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a data detection method and apparatus, an object recommendation method and apparatus, and a device and a medium. The data detection method comprises: acquiring a target sample matrix to be detected, wherein the target sample matrix comprises sample features of sample data; on the basis of position information of the sample features in the target sample matrix, generating first detection data of the target sample matrix; on the basis of feature values of the sample features, generating second detection data of the target sample matrix; and on the basis of the first detection data and the second detection data, detecting whether the target sample matrix is a sparse sample matrix. By means of the embodiments of the present application, accurate identification of the difference between feature data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Data detection method, object recommendation method, device, equipment and medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311706096.9, and invention name “Data Detection Method, Object Recommendation Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence processing technology, and in particular to a data detection method, object recommendation method, device, equipment and medium. Background Art

[0003] With the continuous development of artificial intelligence technology, machine learning has been widely used in various scenarios. In the machine learning process, a matrix is ​​usually used to store the feature data of multiple samples, and this matrix is ​​used as training data for model training. When the number of zero-valued elements in the matrix is ​​much greater than the number of non-zero elements, and the distribution of non-zero elements is irregular, the matrix is ​​considered a sparse matrix. Because sparse matrices represent that the differences between the multiple feature data they store are small, model training based on sparse matrices will increase the learning difficulty of the model and reduce model performance. It can be seen that how to accurately identify the differences between feature data is a technical problem that urgently needs to be solved.

[0004] Summary of the Invention

[0005] The present application provides a data detection method, object recommendation method, device, equipment and medium to accurately identify the differences between feature data.

[0006] In a first aspect, an embodiment of the present application provides a data detection method, comprising:

[0007] Acquire a target sample matrix to be detected; the target sample matrix includes sample features of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0008] generating first detection data of the target sample matrix based on position information of the sample features in the target sample matrix;

[0009] generating second detection data of the target sample matrix based on the eigenvalues ​​of the sample features;

[0010] According to the first detection data and the second detection data, it is detected whether the target sample matrix is ​​a sparse matrix.

[0011] In a second aspect, an embodiment of the present application provides an object recommendation method, comprising:

[0012] Obtain the target user's object operation data and user profile data;

[0013] The object operation data and the user profile data are predicted and processed by an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method described in the first aspect;

[0014] According to the prediction result, a target object to be recommended to the target user is determined.

[0015] In a third aspect, an embodiment of the present application provides a data detection device, including:

[0016] An acquisition module is configured to acquire a target sample matrix to be detected; the target sample matrix includes sample features of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0017] A first generating module, configured to generate first detection data of the target sample matrix based on position information of the sample feature in the target sample matrix;

[0018] A second generating module, configured to generate second detection data of the target sample matrix based on the eigenvalues ​​of the sample features;

[0019] A detection module is configured to detect whether the target sample matrix is ​​a sparse sample matrix based on the first detection data and the second detection data.

[0020] In a fourth aspect, an embodiment of the present application provides an object recommendation device, comprising:

[0021] The acquisition module is used to obtain the target user's object operation data and user portrait data;

[0022] A prediction module, configured to perform prediction processing on the object operation data and the user profile data using an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method described in the first aspect;

[0023] A determination module is used to determine the target object to be recommended to the target user based on the prediction result.

[0024] In a fifth aspect, an embodiment of the present application provides an electronic device, including:

[0025] A processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions are configured to be executed by the processor, the executable instructions including steps for executing the data detection method provided in the first aspect above, or the executable instructions including steps for executing the object recommendation method provided in the second aspect above.

[0026] In a sixth aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, wherein the executable instructions enable a computer to execute the data detection method provided in the first aspect, or the executable instructions include instructions for executing the object recommendation method provided in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate one or more embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0028] FIG1 is a schematic diagram of the composition of a data detection system provided in an embodiment of the present application;

[0029] FIG2 is a schematic diagram of a first flow chart of a data detection method provided in an embodiment of the present application;

[0030] FIG3 is a schematic diagram of a second flow chart of a data detection method provided in an embodiment of the present application;

[0031] FIG4 is a schematic diagram of a third flow chart of a data detection method provided in an embodiment of the present application

[0032] FIG5 is a schematic diagram of a fourth flow chart of a data detection method provided in an embodiment of the present application;

[0033] FIG6 is a fifth flow chart of a data detection method provided in an embodiment of the present application;

[0034] FIG7 is a flow chart of an object recommendation method provided in an embodiment of the present application;

[0035] FIG8 is a schematic diagram of the module composition of a data detection device provided in an embodiment of the present application;

[0036] FIG9 is a schematic diagram of the module composition of an object recommendation device provided in an embodiment of the present application;

[0037] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present application, the technical solutions in one or more embodiments of the present application will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on one or more embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of this document.

[0039] A matrix is ​​considered sparse when the number of zero-valued elements (e.g., elements with zero values) is significantly greater than the number of nonzero elements (e.g., elements with non-zero values), and the distribution of nonzero elements is irregular. Otherwise, the matrix is ​​considered dense. Sparse matrices are used in many scenarios and for many reasons. For example, in e-commerce, there are tens of thousands of products. Since it is impossible for each customer to purchase all of them, a customer's purchase history only covers a small portion of the vast product pool. Therefore, the purchase history of multiple customers can form a sparse matrix. Another example is text mining. To compare articles on the same topic, a group of keywords is often selected and then compared by comparing their frequency of occurrence in each article. This group of keywords may contain tens of thousands of keywords, while each article may only contain tens to hundreds of them. Therefore, the resulting matrix is ​​also a sparse matrix. For example, a sparse matrix may be generated in the word vector representation. For example, an article may contain hundreds or thousands of different words, so the vector representation dimension of each word is hundreds or thousands of dimensions. When using the one-hot method to represent words, the word vector corresponding to each word has only the position of the word as 1, and the rest of the positions are all 0, so the vectors of these words form a sparse matrix.

[0040] Currently, in machine learning, a matrix is ​​typically used to store sample feature data, and this matrix is ​​used as training data for model training and optimization calculations. In practical applications, the sparsity of a matrix characterizes the variability between the feature data within it. Specifically, when a sparse matrix represents a small variability between the multiple feature data within it, the separability between the feature data is unclear. When a dense matrix represents a large variability between the multiple feature data within it, the separability between the feature data is clear. If the training data used for model training is a sparse matrix, the variability between the feature data is small, and the separability between the feature data is unclear, which increases the difficulty of model learning and reduces model performance.

[0041] In order to solve this technical problem, one approach in the prior art is to count the number of zeros contained in the matrix. However, this method cannot compare the sparsity of matrices in different dimensions and ignores the impact of the distribution of other element values. For example, the matrix and matrix Containing the same number of zero-valued elements, according to this method, the sparsity of the two matrices is the same, which is obviously wrong. Another approach in the prior art is to count the proportion of zero-valued elements in the matrix, that is, the number of zero-valued elements divided by the total number of elements in the matrix. However, this method only considers the zero-valued elements and ignores the impact of the distribution of other element values. For example, the matrix and matrix According to this method, the sparsity of the two matrices is the same, which is obviously wrong. Another approach in the prior art is to calculate the ratio of the mean value to the maximum value of the matrix elements, but this method does not take into account the influence of the extreme values ​​of the elements and also ignores the influence of the distribution of the element values. For example, the matrix and matrix According to this method, the sparsity of the two matrices is the same, which is obviously a wrong conclusion. It can be seen that there is currently no effective detection method for matrix sparsity, that is, there is no effective identification method for the differences between feature data. Based on this, an embodiment of the present application provides a data detection method. When a target sample matrix to be detected is obtained, the target sample matrix includes sample features of multiple sample data, the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix, the first detection data of the target sample matrix is ​​generated; and based on the eigenvalues ​​of the multiple sample features included in the target sample matrix, the second detection data of the target sample matrix is ​​generated; and based on the first detection data and the second detection data, whether the target sample matrix is ​​a sparse matrix is ​​detected. Because the first detection data characterizes the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix, and the second detection data characterizes the degree to which the eigenvalues ​​of the multiple sample features included in the target sample matrix tend to zero. Therefore, the data detection method provided by the present application not only takes into account the influence of sample features with zero eigenvalues ​​on the sparsity of the target sample matrix, but also takes into account the influence of sample features with non-zero eigenvalues ​​on the sparsity of the target sample matrix, and can accurately measure the overall complexity of the target sample matrix of each sample feature combination and the overall tendency to zero of the eigenvalues ​​of each sample feature. Accurate detection of the sparsity of the target sample matrix is ​​achieved. Since the sparsity of the target sample matrix characterizes the differences between the sample features in the target sample matrix, accurate identification of the differences between the sample features is achieved. Furthermore, it is possible to detect training data with differences in sample features in terms of machine learning to perform model training to reduce the learning difficulty of the model and improve model performance.

[0042] The data detection method provided in the embodiments of the present application can be implemented by various electronic devices. In some embodiments, it can be implemented by a terminal device alone, that is, the terminal device alone executes the data detection method provided in the embodiments of the application. In other embodiments, it can be implemented in collaboration between the terminal device and the server; for example, the terminal device sends a sample detection request to the server based on the target sample matrix to be detected; the server obtains the target sample matrix to be detected from the received sample detection request, and executes the data detection method based on the target sample matrix.

[0043] The electronic device for executing the data detection method provided in the embodiments of the present application can be various types of terminal devices or server terminals. Among them, the terminal device can be a smart phone, a tablet computer, a portable notebook, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The server terminal can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0044] In some embodiments, the electronic device provided in this application is a terminal device, which implements the data detection method provided in the embodiment of this application by running a computer program. Among them, the computer program can be a native program or software module in the operating system, or it can be an independent application (Application, referred to as APP), or it can be a small program embedded in other applications, or it can be a web application, etc. In short, the above-mentioned computer program can be an application, module or plug-in in any form.

[0045] In other embodiments, the electronic device provided in this application is a server, and the server is a server cluster deployed in the cloud as an example for explanation. The server can use artificial intelligence as a service (AI as a Service, AIaaS) to provide artificial intelligence services to users. Among them, AIaaS will split several common artificial intelligence services and provide independent or packaged services in the cloud. This service model is similar to an AI theme mall. All users can access it through the application programming interface to use one or more artificial intelligence services provided by AIaaS. For example, one of the artificial intelligence services can be a data detection service, that is, the server in the cloud is encapsulated with the data detection program provided in the embodiment of this application. The terminal device responds to the user's sample detection operation by calling the data detection service in the cloud service so that the server deployed in the cloud calls the encapsulated data detection program, generates the first detection data of the target sample matrix based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix; generates the second detection data of the target sample matrix based on the eigenvalues ​​of the multiple sample features included in the target sample matrix; detects whether the target sample matrix is ​​a sparse sample matrix based on the first detection data and the second detection data, and sends the detection result to the terminal device, which displays the received detection result. The terminal device and the server may be connected directly or indirectly via wired or wireless communication methods, and this application does not impose any specific restrictions on this.

[0046] An exemplary data detection system is described by taking a terminal device and a server collaboratively implementing the data detection method provided in an embodiment of the present application as an example. As shown in FIG1 , the system includes: a terminal device, a network, and a server. The terminal device is connected to the server via a network, and the network can be a wide area network, a local area network, or a combination of a wide area network and a local area network. It should be noted that FIG1 illustrates an exemplary embodiment of a mobile phone as the terminal device, but is not limited thereto.

[0047] A client is running in the terminal device, which can be a native program or software module in the operating system of the terminal device, or an independent application (Application, referred to as APP), or a small program embedded in other applications, or a web application, etc. The client responds to the user operation, obtains the target sample matrix to be detected, generates a sample detection request based on the target sample matrix, and sends the sample detection request to the server through the network. The server receives the sample detection request from the terminal device, obtains the target sample matrix to be detected from the sample detection request; based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix, generates the first detection data of the target sample matrix; based on the eigenvalues ​​of the multiple sample features included in the target sample matrix, generates the second detection data of the target sample matrix; based on the first detection data and the second detection data, detects whether the target sample matrix is ​​a sparse sample matrix; and sends the detection result to the terminal device through the network. The terminal device displays the received detection result.

[0048] FIG2 is a flow chart of a data detection method provided by one or more embodiments of the present application. The method in FIG2 can be executed by the above-mentioned electronic device. As shown in FIG2, the method includes the following steps:

[0049] Step S102: obtaining a target sample matrix to be detected; the target sample matrix includes sample features of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0050] In some embodiments, the electronic device may be a terminal device, and a client for performing data detection is provided in the electronic device. The user can operate the client to edit the target sample matrix to be detected and submit a sample detection request; accordingly, obtaining the target sample matrix to be detected in step S102 may include: the electronic device receives the sample detection request submitted by the user, and obtains the target sample matrix to be detected from the sample detection request. In other embodiments, the electronic device may be a server, and the server is connected to the terminal device for communication; accordingly, obtaining the target sample matrix to be detected in step S102 may include: the electronic device may receive the sample detection request sent by the terminal device, and obtain the target sample matrix to be detected from the sample detection request. In some further embodiments, the electronic device may be a terminal device or a server; accordingly, obtaining the target sample matrix to be detected in step S102 may include: the electronic device may obtain a specified type of target data (such as a user's viewing record of each video, etc.) in a specified database based on the received sample detection request, generate an initial matrix according to the target data, and perform dimensionality reduction processing on the initial matrix to obtain the target sample matrix to be detected. The method for obtaining the target sample matrix is ​​not specifically limited in this application, and it can be set as needed in actual applications.

[0051] Furthermore, in some embodiments, the sample data may include voice data, and the sample features include voice features. In one embodiment, the target sample matrix may include voice features of multiple voice data. The voice features may be the number of times the voice data is played, forwarded, or commented on, etc. In other embodiments, the sample data may include video data, and the sample features include video features. In one embodiment, the target sample matrix may include video features of multiple video data. The video features may be the number of times the video data is played, forwarded, or commented on, etc. In yet other embodiments, the sample data may include text data, and the sample features include text features. In one embodiment, the target sample matrix may include text features of multiple text data. The text features may be the number of times the text data is viewed, forwarded, commented on, or the number of times a keyword appears, etc.

[0052] The target sample matrix can be m*n dimensional, that is, the target sample matrix includes m rows and n columns, and has m*n elements. The target sample matrix includes m*n features, the m*n elements correspond one-to-one to the m*n features, and the element values ​​of the m*n elements correspond one-to-one to the eigenvalues ​​of the m*n features; m and n are both integers greater than 1. The specific meanings of the sample features included in the target sample matrix can be set as needed in actual applications. As an example, each row of the sample matrix corresponds to a user and each column corresponds to a voice data, then the element a of the target sample matrix ijThe corresponding voice feature 1 can represent the number of times the user corresponding to the i-th row plays the voice data corresponding to the j-th column; element a ij The element value of the speech feature 1 is the feature value "number of plays", such as element a ij =3, then the eigenvalue of voice feature 1 is 3, that is, the number of times the user corresponding to the i-th row plays the voice data corresponding to the j-th column is 3. As another example, each row of the sample matrix corresponds to a user group and each column corresponds to a video data, then the element a of the target sample matrix is ij The corresponding video feature 1 can represent the number of times the user group corresponding to the i-th row (for example, the user group aged 20-25 years old, etc.) forwards the video data corresponding to the j-th column; element a ij The element value of is the feature value of video feature 1 "number of forwardings", such as element a ij =4, then the feature value of video feature 1 is 4, that is, the number of times the user group corresponding to the i-th row forwarded the video data corresponding to the j-th column is 4. As another example, each row of the target sample matrix corresponds to an article and each column corresponds to a keyword, then the element a of the target sample matrix is ij The corresponding text feature 1 indicates the number of times the keyword corresponding to the jth column appears in the article corresponding to the ith row; element a ij The element value of is the feature value "number of occurrences" of text feature 1, such as element a ij =5, then the feature value of text feature 1 is 5, that is, the keyword corresponding to the jth column appears 5 times in the article corresponding to the i-th row. Among them, 1≤i≤m, 1≤j≤n.

[0053] Step S104 : generating first detection data of the target sample matrix based on the position information of the sample features in the target sample matrix.

[0054] Taking into account the fact that the existing matrix sparsity determination methods all ignore the influence of the distribution of non-zero elements on the matrix sparsity, the accuracy of the matrix sparsity detection is relatively low, and the accuracy of the difference identification between feature data is relatively low. Based on this, in an embodiment of the present application, based on the position information of each feature included in the target sample matrix in the target sample matrix, a first detection data is generated to characterize the complexity of the target sample matrix of the combination of multiple sample features. The first detection data characterizes the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix. Among them, the higher the complexity, the smaller the sparsity of the target sample matrix, and the greater the difference between the sample features in the target sample matrix; conversely, the lower the complexity, the greater the sparsity of the target sample matrix, and the smaller the difference between the sample features in the target sample matrix.

[0055] Since each element in the target sample matrix corresponds to a sample feature, it can be understood that for an m*n dimensional target sample matrix, it is obtained by permuting and combining m*n sample features. Different permutations and combinations of m*n sample features often have different levels of complexity, and accordingly, the sparsity of the target sample matrix obtained is also different. For example, if m=2 and n=3, for a given 6 sample features, the target sample matrix obtained by combining them may include and In the two, the complexity of the arrangement and combination of sample features is different, and the sparsity is also different. In order to accurately measure the sparsity of the target sample matrix, in an embodiment of the present application, the complexity of the arrangement and combination of multiple sample features included in the target sample matrix is ​​characterized by the first detection data. Further, the difference in the arrangement and combination of sample features also characterizes the different positions of the sample features in the target sample matrix, and the position of the sample features in the target sample matrix also reflects the distribution of the sample features in the target sample matrix. Therefore, the first detection data can also characterize the complexity of the distribution of multiple sample features. Further, since the element value of each element in the target sample matrix corresponds to the eigenvalue of a sample feature, the first detection data can also characterize the complexity of the target sample matrix of the eigenvalue combination of multiple sample features, and can also characterize the complexity of the distribution of the eigenvalues ​​of multiple sample features.

[0056] Step S106 : generating second detection data of the target sample matrix based on the eigenvalues ​​of the sample features in the target sample matrix.

[0057] Taking into account that the existing method for determining the sparsity of a matrix only considers the proportion of the number of zero-valued elements, and does not consider the tendency to zero of the matrix as a whole, the accuracy of the determined matrix sparsity is relatively low, that is, the accuracy of identifying the differences between the feature data is relatively low. Based on this, in an embodiment of the present application, based on the eigenvalues ​​of each feature included in the target sample matrix, second detection data is generated to characterize the degree of tendency to zero of the eigenvalues ​​of each sample feature in the target sample matrix. The second detection data characterizes the degree of tendency to zero of the eigenvalues ​​of each sample feature; wherein, the greater the tendency to zero, the more sample features the eigenvalues ​​tend to zero, and the smaller the difference between the sample features; conversely, the smaller the tendency to zero, the more sample features the eigenvalues ​​deviate from zero, and the greater the difference between the sample features.

[0058] Step S108 : detecting whether the target sample matrix is ​​a sparse sample matrix according to the first detection data and the second detection data.

[0059] If the target sample matrix is ​​detected to be a sparse sample matrix, the differences between the sample features in the target sample matrix are determined to be small; if the target sample matrix is ​​detected to be not a sparse sample matrix, the differences between the sample features in the target sample matrix are determined to be large. In other words, detecting the sparsity of the target sample matrix can identify the differences between the sample features in the target sample matrix; and the higher the accuracy of the matrix sparsity detection, the higher the accuracy of the sample feature difference identification.

[0060] In one or more embodiments of the present application, when a target sample matrix to be detected is obtained, first detection data of the target sample matrix is ​​generated based on the position information of the sample features of the multiple sample data included in the target sample matrix in the target sample matrix; and second detection data of the target sample matrix is ​​generated based on the eigenvalues ​​of the multiple sample features included in the target sample matrix; and based on the first detection data and the second detection data, whether the target sample matrix is ​​a sparse matrix is ​​detected. Since the first detection data characterizes the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix, and the second detection data characterizes the degree to which the eigenvalues ​​of the multiple sample features included in the target sample matrix tend to zero, accurate detection of the sparsity of the target sample matrix is ​​achieved. Since the sparsity of the target sample matrix characterizes the differences between the various sample features in the target sample matrix, accurate identification of the differences between the sample features is achieved. Furthermore, in terms of machine learning, it is possible to detect training data with differences in sample features for model training to reduce the learning difficulty of the model and improve model performance.

[0061] In order to accurately measure the complexity of the distribution of multiple features included in the target sample matrix, in one or more embodiments of the present application, the first detection data is generated based on the feature distribution information corresponding to each target vector. Specifically, as shown in Figure 3, step S104 may include the following steps S104-2 to S104-6:

[0062] Step S104 - 2 : determining the sample feature combination corresponding to each target vector in the target sample matrix according to the position information of the sample features in the target sample matrix, wherein the target vector includes a row vector or a column vector.

[0063] In some embodiments, the target vector may be a row vector, and correspondingly, the position information may be a row number. In other embodiments, the target vector may be a column vector, and correspondingly, the position information may be a column number. Specifically, when the target vector is a row vector, the sample feature combination corresponding to each target vector may be determined based on the row number of each sample feature in the multiple sample features included in the target sample matrix. When the target vector is a column vector, the sample feature combination corresponding to each target vector may be determined based on the column number of each sample feature in the multiple sample features included in the target sample matrix. The following descriptions are all based on the target vector being a row vector. When the target vector is a column vector, reference may be made to the relevant description of the target vector being a row vector, and the rows in the relevant description may be replaced with columns.

[0064] As an example, the target sample matrix is If the target vector is a row vector, the target sample matrix includes two row vectors, which are denoted as target vector 1 and target vector 2 in order from top to bottom. The row numbers of the sample features corresponding to eigenvalues ​​0, 5, and 6 are all 1, so the sample feature combination corresponding to target vector 1 is determined to be the sample features corresponding to eigenvalues ​​0, 5, and 6, that is, the sample feature combination corresponding to target vector 1 is the sample features corresponding to the first row eigenvalues; the row numbers of the sample features corresponding to eigenvalues ​​7, 0, and 0 are all 2, so the sample feature combination corresponding to target vector 2 is determined to be the sample features corresponding to eigenvalues ​​7, 0, and 0, that is, the sample feature combination corresponding to target vector 2 is the sample features corresponding to the second row eigenvalues. Similarly, when the target vector is a column vector, the target sample matrix includes target vector 1, target vector 2, and target vector 3 in order from left to right; the sample feature combination corresponding to target vector 1 is the sample features corresponding to the eigenvalues ​​0 and 7 in the first column, the sample feature combination corresponding to target vector 2 is the sample features corresponding to the eigenvalues ​​5 and 0 in the second column, and the sample feature combination corresponding to target vector 3 is the sample features corresponding to the eigenvalues ​​6 and 0 in the third column.

[0065] Step S104 - 4 , determining sample feature distribution information of each sample feature combination.

[0066] Specifically, for each sample feature combination, the eigenvalues ​​of the sample feature combination are deduplicated to obtain a first number of target eigenvalues; wherein the target eigenvalue is the eigenvalue remaining after the eigenvalues ​​of the sample feature combination are deduplicated; for each target eigenvalue, a target proportion of the number of times the target eigenvalue appears in the eigenvalues ​​of the sample feature combination is determined; the first number and the target proportion are determined as the sample feature distribution information of the sample feature combination.

[0067] More specifically, for each sample feature combination, the eigenvalues ​​of the sample feature combination are deduplicated to obtain the target eigenvalues ​​after deduplication, and a first number of the target eigenvalues ​​is counted; and, for each target eigenvalue, the number of occurrences of the target eigenvalue in the eigenvalues ​​of the corresponding sample feature combination is identified, and based on the number of occurrences and the total number of eigenvalues ​​of the corresponding sample feature combination, a target proportion of the number of occurrences of the target eigenvalue in the eigenvalues ​​of the corresponding sample feature combination is determined; and the first number and the determined target proportions are determined as the sample feature distribution information of the corresponding sample feature combination.

[0068] As an example, the target vector is a row vector, and a certain sample feature combination includes 9 sample features, whose eigenvalues ​​are 0, 0, 1, 2, 0, 3, 1, 0, 1 in order from left to right. The remaining target eigenvalues ​​after deduplication are 0, 1, 2, 3, that is, the first number is 4; the target proportion of the number of occurrences of the target eigenvalue 0 is 4 / 9, the target proportion of the number of occurrences of the target eigenvalue 1 is 3 / 9, the target proportion of the number of occurrences of the target eigenvalue 2 is 1 / 9, and the target proportion of the number of occurrences of the target eigenvalue 3 is 1 / 9.

[0069] Step S104 - 6 : generating first detection data of the target sample matrix according to the sample feature distribution information.

[0070] Specifically, the sample feature complexity of each target vector is determined based on the sample feature distribution information. The sample feature complexity represents the complexity of the permutations and combinations of the multiple sample features corresponding to the target vector to form the target vector. First detection data of the target sample matrix is ​​generated based on the sample feature complexity and the second number of target vectors. The first detection data represents the complexity of the permutations and combinations of the multiple sample features included in the target sample matrix.

[0071] More specifically, the sample feature complexity of the corresponding target vector is determined based on the first quantity and target proportion in each sample feature distribution information; and the sum of the sample feature complexities of each target vector is determined, and the average value of the feature complexity of each target vector is determined based on the sum of the sample feature complexities of each target vector and the second quantity of the target vector; the average value is determined as the first detection data of the target sample matrix.

[0072] In some implementations, the first detection data is recorded as ω, and then:

[0073] The target vector is used as the row vector for explanation, where a i Indicates that the first number of target eigenvalues ​​in the eigenvalues ​​of the sample feature combination corresponding to the i-th target vector (i.e., the i-th row) is a; trepresents the tth target eigenvalue in the target eigenvalue of the sample feature combination corresponding to the i-th target vector; p(x t ) represents the target proportion of the number of occurrences of the t-th target eigenvalue in the target eigenvalue of the sample feature combination corresponding to the i-th target vector; log2p(x t ) represents the base 2 p(x t ) represents the sample feature complexity of the i-th target vector; represents the sum of the sample feature complexities of each target vector; m is the second number of target vectors (i.e., the total number of rows in the target sample matrix); ω represents the average value of the sample feature complexity of each target vector, i.e., the first detection data; the larger ω is, the higher the complexity of the arrangement and combination of multiple sample features in the target sample matrix, and the smaller the sparsity of the target sample matrix.

[0074] It can be seen that when generating the first detection data, it is based on the position information of each sample feature included in the target sample matrix, which involves both features with zero eigenvalues ​​and features with non-zero eigenvalues. Therefore, the sparsity of the target sample matrix is ​​detected from the perspective of the overall distribution of each sample feature in the target sample matrix, rather than only detecting the sparsity of the target sample matrix based on features with zero eigenvalues. Therefore, the accuracy of matrix sparsity detection is greatly improved, and the accuracy of difference identification between multiple sample data corresponding to the matrix is ​​improved.

[0075] While detecting the sparsity of the target sample matrix from the perspective of overall distribution, the present application also detects the sparsity of the target sample matrix from the perspective of the overall null tendency of the target sample matrix in terms of numerical value, so as to further ensure the accuracy of data difference identification. Specifically, as shown in Figure 4, step S106 may include the following steps S106-2 and S106-4:

[0076] Step S106-2: Determine the characteristic zeroing degree of each target vector according to the characteristic value of the sample characteristic in the sample characteristic combination.

[0077] Specifically, for each target vector, the geometric length of the target vector is determined based on the eigenvalues ​​of each sample feature in the sample feature combination corresponding to the target vector. The characteristic zeroing degree of each target vector is determined based on the eigenvalues ​​and geometric lengths of each sample feature in the sample feature combination corresponding to the target vector. The characteristic zeroing degree indicates the degree to which the eigenvalues ​​of the multiple sample features corresponding to the target vector tend to zero.

[0078] More specifically, for each target vector, the eigenvalue of each sample feature in the sample feature combination corresponding to the target vector is squared to obtain a first value; the first values ​​are summed and then squared to obtain the geometric length of the target vector; the absolute value of the eigenvalue of each sample feature in the sample feature combination corresponding to the target vector is taken and then added to obtain a second value, and the second value is divided by the geometric length to obtain the characteristic zeroing degree of each target vector.

[0079] Step S106 - 4 : generating second detection data of the target sample matrix according to the characteristic zeroing degree, the second number of target vectors, and the third number of sample features corresponding to the target vectors.

[0080] In some embodiments, the second detection data is recorded as γ, and then:

[0081] The target vector is used as the row vector for explanation, |y i,j | represents the absolute value of the eigenvalue of the jth sample feature in the sample feature set corresponding to the i-th target vector in the target sample matrix (that is, the absolute value of the j-th element value in the i-th row), y i,j 2 Represents the square of the eigenvalue of the jth sample feature in the sample feature set corresponding to the i-th target vector (i.e., the square of the jth element value in the i-th row); represents the geometric length of the i-th target vector; m is the second number of target vectors (i.e., the total number of rows in the target sample matrix), and n is the third number of sample features corresponding to each target vector (i.e., the total number of columns in the target sample matrix).

[0082] It can be seen that in the process of generating the second detection data, it is based on the eigenvalue of each sample feature included in the target sample matrix, involving both zero eigenvalues ​​and non-zero eigenvalues. Therefore, the sparsity of the target sample matrix is ​​detected from the overall numerical value of each sample feature in the target sample matrix, rather than only based on the zero eigenvalue to detect the sparsity of the target sample matrix, thereby greatly improving the accuracy of matrix sparsity detection. Furthermore, since the second detection data is generated based on the geometric length of the target vector, and the geometric length of the target vector effectively avoids the influence of extreme values ​​(deviations from normal values) in the existing methods, the accuracy of matrix sparsity detection is further improved, that is, the accuracy of difference recognition between sample data is improved.

[0083] After generating the first detection data and the second detection data, it is possible to detect whether the target sample matrix is ​​a sparse matrix based on the first detection data and the second detection data. Specifically, as shown in FIG5 , step S108 may include the following steps S108-2 to S108-6:

[0084] Step S108-2, comparing the first detection data with a first threshold to obtain a first comparison result;

[0085] Step S108-4, comparing the second detection data with the second threshold to obtain a second comparison result;

[0086] Step S108-6: Determine whether the target sample matrix is ​​a sparse matrix based on the first comparison result and the second comparison result.

[0087] Specifically, if the first comparison result indicates that the first test data is greater than the first threshold, and the second comparison result indicates that the second test data is greater than the second threshold, then it is determined that the target sample matrix is ​​not a sparse matrix; otherwise, it is determined that the target sample matrix is ​​a sparse matrix. The first threshold and the second threshold can be obtained by pre-testing the training data in each matrix form based on the above method, obtaining the first test data and the second test data of each training data, performing model training based on the training data, and determining the first test data corresponding to the target model when the target model is obtained as the first threshold, and determining the second test data corresponding to the target model when the target model is obtained as the second threshold.

[0088] Therefore, the sparsity of the target sample matrix is ​​detected together with the overall distribution of each feature in the target sample matrix and the overall tendency of the eigenvalues ​​of each feature to zero, which improves the accuracy of the detection results and realizes the effective detection of matrix sparsity, that is, the effective identification of differences between sample data.

[0089] As mentioned above, the above-mentioned matrix sparsity detection device can be applied to detect training data in a matrix form. Therefore, in one or more embodiments of the present application, the following step S110 may be further included after step S108:

[0090] Step S110: If the detection result indicates that the target sample matrix is ​​not a sparse matrix, iterative training is performed on the network to be trained using the target sample matrix.

[0091] Specifically, if the detection result indicates that the target sample matrix is ​​not a sparse matrix, the target sample matrix is ​​determined as target training data, and when it is determined that the training condition is met, the target training data is used to iteratively train the network to be trained to obtain a target model. Determining that the training condition is met may include: determining that the training condition is met if it is determined that the amount of target training data reaches a preset amount.

[0092] Therefore, before training, the sparsity of the training data in each matrix form is first detected and sorted, and training is performed based on the target training data that is not a sparse matrix, which ensures that the features in the target training data are separable, thereby greatly reducing the learning difficulty of the model and improving the model performance.

[0093] Furthermore, considering that there may be a large number of sample matrices in the machine learning process, it is often necessary to compare the sparsity of the large number of sample matrices to obtain target training data. Based on this, in one or more embodiments of the present application, step S108 may further include the following steps S112 and S114:

[0094] Step S112: determining the sparsity parameter of the target sample matrix according to the first detection data and the second detection data.

[0095] Specifically, let the sparsity parameter be θ, then: Here, ω represents the first detection data, and γ represents the second detection data.

[0096] Step S114 , performing sparsity comparison processing on the target sample matrix and the matrix to be compared according to the sparsity parameter.

[0097] Specifically, for each matrix to be compared, the sparsity parameter of the matrix to be compared is determined; the sparsity parameter of the target sample matrix is ​​compared with the sparsity parameter of the matrix to be compared; if the comparison result indicates that the sparsity parameter of the target sample matrix is ​​greater than the sparsity parameter of the matrix to be compared, then the sparsity of the target sample matrix is ​​determined to be less than the sparsity of the matrix to be compared; if the comparison result indicates that the sparsity parameter of the target sample matrix is ​​less than the sparsity parameter of the matrix to be compared, then the sparsity of the target sample matrix is ​​determined to be greater than the sparsity of the matrix to be compared. In this way, based on the first test data and the second test data, the sparsity comparison between different sample matrices is achieved.

[0098] In one or more embodiments of the present application, when a target sample matrix to be detected is obtained, the target sample matrix includes sample features of multiple sample data, the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix, first detection data of the target sample matrix is ​​generated; and based on the eigenvalues ​​of the multiple sample features included in the target sample matrix, second detection data of the target sample matrix is ​​generated; and based on the first detection data and the second detection data, whether the target sample matrix is ​​a sparse matrix is ​​detected. Since the first detection data represents the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix, and the second detection data represents the degree to which the eigenvalues ​​of the multiple sample features included in the target sample matrix tend to zero; therefore, not only the influence of the sample features with zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, but also the influence of the sample features with non-zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, and the overall complexity of the target sample matrix of each sample feature combination and the overall degree to which the eigenvalues ​​of each sample feature tend to zero can be accurately measured, thereby achieving accurate detection of the sparsity of the target sample matrix. Because the sparsity of the target sample matrix characterizes the differences between the features of each sample in the target sample matrix, it can accurately identify the differences between sample features. Furthermore, in machine learning, it can detect training data with differences in sample features for model training, thereby reducing the learning difficulty of the model and improving model performance.

[0099] In some specific embodiments, aforementioned sample data can be video data, aforementioned sample features can be video features, aforementioned target sample matrix can be target video sample matrix, and aforementioned target model is video recommendation model.Because video may produce hundreds of millions every day, and each user objectively can not watch all videos once, so user's video operation data (such as viewing record, forwarding record, comment record etc.) is for a small part of video in all videos, that is, user's video operation data can constitute an initial sample matrix, and this initial sample matrix is ​​a sparse matrix.If this initial sample matrix is ​​used to carry out model training now, because most of features therein are zero, there is no separability of feature, that is, the difference between features is smaller, therefore must increase model learning difficulty, can not obtain a better video recommendation model, and then affect the video recommendation effect to user.Therefore, data processing can be carried out by above-mentioned data processing mode. That is, a target video sample matrix to be detected is obtained, wherein the target video sample matrix includes video features of multiple video data; based on the position information of each video feature in the target video sample matrix, first detection data of the target video sample matrix is ​​generated; the first detection data represents the complexity of the target video sample matrix composed of each video feature; and based on the eigenvalue of each video feature, second detection data of the target video sample matrix is ​​generated; the second detection data represents the degree to which the eigenvalue of each video feature tends to zero; and based on the first detection data and the second detection data, whether the target video sample matrix is ​​a sparse sample matrix is ​​detected. Specifically, as shown in FIG6 , the following steps S10 to S26 may be included:

[0100] Step S10: obtaining a target video sample matrix to be detected, wherein the target video sample matrix includes video features of a plurality of video data.

[0101] In some implementations, video operation data for multiple users can be obtained from a designated database. Each row corresponds to a user, and each column corresponds to a video data item. An initial sample matrix is ​​constructed based on the obtained video operation data. Dimensionality reduction is then performed on this initial sample matrix to obtain a target video sample matrix to be tested. The video operation data can include at least one of playback history, forwarding history, comment history, favorite history, and playback duration. For example, dimensionality reduction can be performed on an 8000*100000 initial sample matrix to obtain a 200*1000 target sample matrix.

[0102] It should be noted that the method for obtaining the target video sample matrix is ​​not limited to the above method, and it can be set as needed in actual applications.

[0103] Step S12: determining the video feature combination corresponding to each target vector in the target video sample matrix according to the position information of the video features in the target video sample matrix, wherein the target vector includes a row vector or a column vector.

[0104] Step S14: determining video feature distribution information of each video feature combination.

[0105] Specifically, for each video feature combination, the eigenvalues ​​of the video feature combination are deduplicated to obtain a first number of target eigenvalues; the target eigenvalues ​​are the eigenvalues ​​remaining after the eigenvalues ​​of the video feature combination are deduplicated; for each target eigenvalue, a target proportion of the number of times the target eigenvalue appears in the eigenvalues ​​of the video feature combination is determined; the first number and the target proportion are determined as the sample video feature distribution information of the video feature combination.

[0106] Step S16, determining the video feature complexity of each corresponding target vector based on each video feature distribution information; the sample feature complexity represents the complexity of the target vector combined with multiple video features corresponding to the target vector.

[0107] Step S18: Generate first detection data of the target video sample matrix according to the complexity of each video feature and the second number of target vectors.

[0108] Step S20 , determining the characteristic zeroing degree of each target vector according to the characteristic values ​​of each video feature in the video feature combination; the characteristic zeroing degree represents the degree to which the characteristic values ​​of the multiple video features corresponding to the target vector tend to zero.

[0109] Step S22 : generating second detection data of the target video sample matrix according to the feature zeroing degree, the second number of the target vectors, and the third number of the video features corresponding to the target vectors.

[0110] Step S24 : detecting whether the target video sample matrix is ​​a sparse sample matrix according to the first detection data and the second detection data.

[0111] Step S26: If the target video sample matrix is ​​not a sparse sample matrix, the target video sample matrix is ​​determined as target training data, and when it is determined that the training conditions are met, the network to be trained is iteratively trained based on the target training data to obtain a video recommendation model.

[0112] Wherein, determining that the training conditions are met may include: if it is determined that the number of target training data reaches a preset number, then determining that the training conditions are met. Furthermore, considering that users of different age groups or users of different occupations are often interested in different video data, based on this, in one or more embodiments of the present application, determining the target video sample matrix as the target training data in step S26 may include: determining the target sample matrix and the user portrait data of each user corresponding to the target sample matrix as the target training data. Wherein, the user portrait data may include at least one of the user's age, occupation, gender, etc. The specific process of iterative training the network to be trained based on the target training data can refer to the relevant technology, which will not be described in detail in this application.

[0113] It should be noted that the specific implementation of steps S10 to S26 can refer to the relevant description above, and the repeated parts will not be repeated here.

[0114] Thus, when the target video sample matrix to be detected is obtained, first detection data of the target video sample matrix is ​​generated based on the position information of the multiple video features included in the target video sample matrix in the target video sample matrix; and second detection data of the target video sample matrix is ​​generated based on the eigenvalues ​​of the multiple video features included in the target video sample matrix; and whether the target video sample matrix is ​​a sparse matrix is ​​detected based on the first detection data and the second detection data. Since the first detection data represents the complexity of the arrangement and combination of the multiple video features included in the target video sample matrix; and the second detection data represents the degree to which the eigenvalues ​​of the multiple video features included in the target video sample matrix tend to zero; therefore, not only the influence of video features with zero eigenvalues ​​on the sparsity of the target video sample matrix is ​​taken into account, but also the influence of video features with non-zero eigenvalues ​​on the sparsity of the target video sample matrix is ​​taken into account, and the overall complexity of the target video sample matrix of each video feature combination and the overall degree to which the eigenvalues ​​of each video feature tend to zero are accurately measured, thereby achieving accurate detection of the sparsity of the target video sample matrix. Since the sparsity of the target video sample matrix represents the differences between the video features in the target video sample matrix, accurate identification of the differences between video features is achieved. Furthermore, in terms of machine learning, it is possible to detect training data with differences in video features for model training, so as to reduce the learning difficulty of the model and improve the model performance.

[0115] Corresponding to the data detection method described above, based on the same technical concept, one or more embodiments of the present application also provide an object recommendation method. FIG7 is a flow chart of an object recommendation method provided by one or more embodiments of the present application. As shown in FIG7, the method includes:

[0116] Step S202: Obtain the target user's object operation data and user portrait data;

[0117] The object can be voice, text, video, etc. When the object is voice or video, the object operation data may include at least one of the following: playback history, forwarding history, comment history, favorite history, and playback duration; when the object is text, the object operation data may include at least one of the following: viewing history, forwarding history, comment history, favorite history, etc. User profile data may include at least one of the user's age, gender, occupation, etc.

[0118] Step S204: Predicting the object operation data and the user profile data using an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iteratively training non-sparse target training data; the target training data is obtained by detecting a plurality of target sample matrices to be detected using the data detection method provided in any of the aforementioned embodiments;

[0119] Specifically, the acquired object operation data and user profile data are input into the object recommendation model for prediction processing to obtain a prediction result. The specific process of the prediction processing can be referenced in related art and will not be described in detail in this application. The prediction result may include the target user's preference score for each preset object, which can be a decimal between 0 and 1.

[0120] Step S206: Determine the target object to be recommended to the target user based on the prediction result.

[0121] Specifically, the preset object corresponding to the highest preference score in the prediction result is determined as the target object to be recommended to the target user.

[0122] In one or more embodiments of the present application, when the object operation data and user portrait data of the target user are obtained, the obtained object operation data and user portrait data are predicted and processed by the object recommendation model to obtain a prediction result; and the target object to be recommended to the target user is determined based on the prediction result. Since the training data used by the object recommendation model used in the prediction process is non-sparse target training data obtained by detecting and processing multiple target sample matrices to be detected based on the aforementioned data detection method, that is, the differences between the features of each sample in the target training data are large, the object recommendation model can better and more easily learn the differences between the features of each sample during the training process, thereby improving the model performance; and then performing prediction processing based on the high-performance object recommendation model improves the accuracy of the prediction results, thereby improving the accuracy of the object recommendation.

[0123] Corresponding to the data detection method described above, based on the same technical concept, one or more embodiments of the present application also provide a data detection device. FIG8 is a schematic diagram of the module composition of a data detection device provided by one or more embodiments of the present application. As shown in FIG8, the device includes:

[0124] An acquisition module 301 is configured to acquire a target sample matrix to be detected; the target sample matrix includes sample features of a plurality of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0125] A first generating module 302 is configured to generate first detection data of the target sample matrix based on position information of the sample feature in the target sample matrix;

[0126] A second generating module 303 is configured to generate second detection data of the target sample matrix based on the eigenvalues ​​of the sample features;

[0127] The detection module 304 is configured to detect whether the target sample matrix is ​​a sparse sample matrix based on the first detection data and the second detection data.

[0128] The data detection device provided by the embodiment of the present application, when obtaining a target sample matrix to be detected, the target sample matrix includes sample features of multiple sample data, the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix, first detection data of the target sample matrix is ​​generated; and based on the eigenvalues ​​of the multiple sample features included in the target sample matrix, second detection data of the target sample matrix is ​​generated; and based on the first detection data and the second detection data, whether the target sample matrix is ​​a sparse matrix is ​​detected. Since the first detection data represents the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix, and the second detection data represents the degree of zeroing of the eigenvalues ​​of the multiple sample features included in the target sample matrix; therefore, not only the influence of the sample features with zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, but also the influence of the sample features with non-zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, and the overall complexity of the target sample matrix of each sample feature combination and the overall degree of zeroing of the eigenvalues ​​of each sample feature can be accurately measured, thereby achieving accurate detection of the sparsity of the target sample matrix. Because the sparsity of the target sample matrix characterizes the differences between the features of each sample in the target sample matrix, it can accurately identify the differences between sample features. Furthermore, in machine learning, it can detect training data with differences in sample features for model training, thereby reducing the learning difficulty of the model and improving model performance.

[0129] It should be noted that the embodiment of the data detection device in this application and the embodiment of the data detection method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding data detection method mentioned above, and the repeated parts will not be repeated.

[0130] Each module in the above-mentioned data detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the terminal device or the server in hardware form, or can be stored in the memory in the terminal device or the server in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0131] Furthermore, corresponding to the object recommendation method described above, based on the same technical concept, one or more embodiments of the present application also provide an object recommendation device. FIG9 is a schematic diagram of the module composition of an object recommendation device provided by one or more embodiments of the present application. As shown in FIG9 , the device includes:

[0132] Acquisition module 401, used to acquire the target user's object operation data and user portrait data;

[0133] Prediction module 402, configured to perform prediction processing on the object operation data and the user profile data using an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method described in the first aspect;

[0134] The determination module 403 is configured to determine the target object to be recommended to the target user based on the prediction result.

[0135] The object recommendation device provided in the embodiment of the present application, when obtaining the object operation data and user portrait data of the target user, performs prediction processing on the obtained object operation data and user portrait data through the object recommendation model to obtain a prediction result; and determines the target object to be recommended to the target user based on the prediction result. Since the training data used by the object recommendation model used in the prediction processing during the training process is non-sparse target training data obtained by detecting and processing multiple target sample matrices to be detected based on the data detection method provided in the first aspect, that is, the differences between the features of each sample in the target training data are large, the object recommendation model can better and more easily learn the differences between the features of each sample during the training process, thereby improving the model performance; and then performing prediction processing based on the high-performance object recommendation model improves the accuracy of the prediction results, thereby improving the accuracy of object recommendation.

[0136] It should be noted that the embodiment of the object recommendation device in this application and the embodiment of the object recommendation method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding object recommendation method mentioned above, and the repeated parts will not be repeated.

[0137] Each module in the object recommendation device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the terminal device or the server in hardware form, or can be stored in the memory in the terminal device or the server in software form, so that the processor can call and execute the corresponding operations of each module.

[0138] Furthermore, corresponding to the data detection method and object recommendation method described above, based on the same technical concept, one or more embodiments of the present application also provide an electronic device, which is used to execute the above-mentioned data detection method and object recommendation method. Figure 10 is a structural schematic diagram of an electronic device provided by one or more embodiments of the present application.

[0139] As shown in Figure 10, electronic devices may have relatively large differences due to different configurations or performances, and may include one or more processors 501 and memory 502, and the memory 502 may store one or more storage applications or data. Among them, the memory 502 can be a temporary storage or a persistent storage. The application stored in the memory 502 may include one or more modules (not shown in the figure), each module may include a series of computer executable instructions in the electronic device. Furthermore, the processor 501 can be configured to communicate with the memory 502 to execute a series of computer executable instructions in the memory 502 on the electronic device. The electronic device may also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input and output interfaces 505, one or more keyboards 506, etc.

[0140] In a specific embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the electronic device, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following:

[0141] Acquire a target sample matrix to be detected; the target sample matrix includes sample features of a plurality of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0142] generating first detection data of the target sample matrix based on position information of the sample features in the target sample matrix;

[0143] generating second detection data of the target sample matrix based on the eigenvalues ​​of the sample features;

[0144] According to the first detection data and the second detection data, it is detected whether the target sample matrix is ​​a sparse matrix.

[0145] One or more embodiments of the present application provide an electronic device that, upon acquiring a target sample matrix to be detected, includes sample features of multiple sample data, the sample data including voice data, video data, or text data, and the sample features including voice features, video features, or text features; generates first detection data of the target sample matrix based on position information of the multiple sample features included in the target sample matrix; and generates second detection data of the target sample matrix based on eigenvalues ​​of the multiple sample features included in the target sample matrix; and detects whether the target sample matrix is ​​a sparse matrix based on the first detection data and the second detection data. Since the first detection data represents the complexity of the arrangement and combination of the multiple sample features included in the target sample matrix, and the second detection data represents the degree to which the eigenvalues ​​of the multiple sample features included in the target sample matrix tend to zero; therefore, not only the influence of sample features with zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, but also the influence of sample features with non-zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, and the overall complexity of the target sample matrix of each sample feature combination and the overall degree to which the eigenvalues ​​of each sample feature tend to zero are accurately measured, thereby achieving accurate detection of the sparsity of the target sample matrix. Because the sparsity of the target sample matrix characterizes the differences between the features of each sample in the target sample matrix, it can accurately identify the differences between sample features. Furthermore, in machine learning, it can detect training data with differences in sample features for model training, thereby reducing the learning difficulty of the model and improving model performance.

[0146] In another specific embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the electronic device, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following:

[0147] Obtain the target user's object operation data and user profile data;

[0148] The object operation data and the user profile data are predicted and processed by an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method described in the first aspect;

[0149] According to the prediction result, a target object to be recommended to the target user is determined.

[0150] The electronic device provided in the embodiment of the present application, when acquiring the object operation data and user portrait data of the target user, performs prediction processing on the acquired object operation data and user portrait data through the object recommendation model to obtain a prediction result; and determines the target object to be recommended to the target user based on the prediction result. Since the training data used by the object recommendation model used in the prediction processing during the training process is non-sparse target training data obtained by detecting and processing multiple target sample matrices to be detected based on the data detection method provided in the first aspect, that is, the differences between the features of each sample in the target training data are large, the object recommendation model can better and more easily learn the differences between the features of each sample during the training process, thereby improving the model performance; and then performing prediction processing based on the high-performance object recommendation model improves the accuracy of the prediction results, thereby improving the accuracy of the object recommendation.

[0151] It should be noted that the embodiments of the electronic device in this application and the embodiments of the data detection method and the image denoising method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding data detection method and image denoising method mentioned above, and the repeated parts will not be repeated.

[0152] Furthermore, corresponding to the data detection method described above, based on the same technical concept, one or more embodiments of the present application further provide a storage medium for storing computer-executable instructions. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer-executable instructions stored in the storage medium are executed by a processor, the following process can be implemented:

[0153] Acquire a target sample matrix to be detected; the target sample matrix includes sample features of a plurality of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features;

[0154] generating first detection data of the target sample matrix based on position information of the sample features in the target sample matrix;

[0155] generating second detection data of the target sample matrix based on the eigenvalues ​​of the sample features;

[0156] According to the first detection data and the second detection data, it is detected whether the target sample matrix is ​​a sparse matrix.

[0157] When the computer-executable instructions stored in the storage medium provided by one or more embodiments of the present application are executed by a processor, when a target sample matrix to be detected is obtained, the target sample matrix includes sample features of multiple sample data, the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; based on the position information of the multiple sample features included in the target sample matrix in the target sample matrix, first detection data of the target sample matrix is ​​generated; and based on the eigenvalues ​​of the multiple sample features included in the target sample matrix, second detection data of the target sample matrix is ​​generated; and based on the first detection data and the second detection data, whether the target sample matrix is ​​a sparse matrix is ​​detected. Since the first detection data characterizes the complexity of the arrangement and combination of multiple sample features included in the target sample matrix, and the second detection data characterizes the degree to which the eigenvalues ​​of the multiple sample features included in the target sample matrix tend to zero; therefore, not only the influence of sample features with zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, but also the influence of sample features with non-zero eigenvalues ​​on the sparsity of the target sample matrix is ​​taken into account, and the overall complexity of the target sample matrix of each sample feature combination and the overall degree to which the eigenvalues ​​of each sample feature tend to zero can be accurately measured, thereby achieving accurate detection of the sparsity of the target sample matrix. Since the sparsity of the target sample matrix characterizes the differences between the sample features in the target sample matrix, accurate identification of the differences between the sample features is achieved. Furthermore, in terms of machine learning, it is possible to detect training data with differences in sample features for model training to reduce the learning difficulty of the model and improve model performance.

[0158] In another specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc., and the computer executable instructions stored in the storage medium, when executed by the processor, can implement the following process:

[0159] Obtain the target user's object operation data and user profile data;

[0160] The object operation data and the user profile data are predicted and processed by an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method described in the first aspect;

[0161] According to the prediction result, a target object to be recommended to the target user is determined.

[0162] When the computer executable instructions stored in the storage medium provided by one or more embodiments of the present application are executed by the processor, when the object operation data and user portrait data of the target user are obtained, the object operation data and user portrait data obtained are predicted and processed by the object recommendation model to obtain a prediction result; and the target object to be recommended to the target user is determined based on the prediction result. Since the training data used by the object recommendation model used in the prediction processing during the training process is non-sparse target training data obtained by detecting and processing multiple target sample matrices to be detected based on the data detection method provided in the first aspect, that is, the differences between the features of each sample in the target training data are large, the object recommendation model can better and more easily learn the differences between the features of each sample during the training process, thereby improving the model performance; and then performing prediction processing based on the high-performance object recommendation model improves the accuracy of the prediction results, thereby improving the accuracy of the object recommendation.

[0163] It should be noted that the embodiment of the storage medium in this application and the embodiment of the data detection method and the image denoising method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding data detection method and image denoising method mentioned above, and the repeated parts will not be repeated.

[0164] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0165] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0166] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0167] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0168] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0169] Those skilled in the art will appreciate that one or more embodiments of the present application may be provided as a method, system, or computer program product. Therefore, one or more embodiments of the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0171] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0173] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0174] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0175] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0176] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0177] One or more embodiments of the present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0178] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.

[0179] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.

Claims

1. A data detection method, the method comprising: Acquire a target sample matrix to be detected; the target sample matrix includes sample features of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; Based on the position information of the sample features in the target sample matrix, generating first detection data of the target sample matrix; generating second detection data of the target sample matrix based on the eigenvalues ​​of the sample features; According to the first detection data and the second detection data, it is detected whether the target sample matrix is ​​a sparse sample matrix.

2. The method according to claim 1, wherein generating first detection data of the target sample matrix based on position information of the sample features in the target sample matrix comprises: Determine, according to the position information of the sample features in the target sample matrix, a sample feature combination corresponding to each target vector in the target sample matrix; Determining sample feature distribution information of each of the sample feature combinations; Generate first detection data of the target sample matrix according to the sample feature distribution information.

3. The method according to claim 2, wherein determining the sample feature distribution information of each sample feature combination comprises: For each of the sample feature combinations, performing deduplication processing on the feature values ​​of the sample feature combination to obtain a first number of target feature values; The target feature value is the feature value remaining after deduplication of the feature value of the sample feature combination; For each of the target feature values, determining a target proportion of the number of times the target feature value appears in the feature values ​​of the sample feature combination; The first quantity and the target proportion are determined as sample feature distribution information of the sample feature combination.

4. The method according to claim 2, wherein generating first detection data of the target sample matrix according to the sample feature distribution information comprises: Determining the sample feature complexity of each of the target vectors according to the sample feature distribution information; First detection data of the target sample matrix is ​​generated according to the sample feature complexity and the second number of the target vectors.

5. The method according to claim 2, wherein generating the second detection data of the target sample matrix based on the eigenvalue of the sample feature comprises: Determining the characteristic zeroing degree of each of the target vectors according to the characteristic values ​​of the sample characteristics in the sample characteristic combination; Second detection data of the target sample matrix is ​​generated according to the characteristic zeroing degree, the second number of the target vectors and the third number of sample features corresponding to the target vectors.

6. The method according to claim 5, wherein determining the characteristic zeroing degree of each target vector according to the characteristic value of the sample characteristic in the sample characteristic combination comprises: For each of the target vectors, determining the geometric length of the target vector according to the characteristic value of the sample feature in the sample feature combination corresponding to the target vector; According to the characteristic value of the sample feature in the sample feature combination corresponding to the target vector and the geometric length, Determine the characteristic zeroing degree of each of the target vectors.

7. The method according to claim 1, wherein detecting whether the target sample matrix is ​​a sparse sample matrix according to the first detection data and the second detection data comprises: Comparing the first detection data with a first threshold to obtain a first comparison result; Comparing the second detection data with a second threshold value to obtain a second comparison result; According to the first comparison result and the second comparison result, it is determined whether the target sample matrix is ​​a sparse sample matrix.

8. The method according to claim 1, after detecting whether the target sample matrix is ​​a sparse sample matrix, the method further comprises: Determining a sparsity parameter of the target sample matrix according to the first detection data and the second detection data; According to the sparsity parameter, a sparsity comparison process is performed on the target sample matrix and the sample matrix to be compared.

9. An object recommendation method, comprising: Obtain the target user's object operation data and user portrait data; The object operation data and the user portrait data are predicted and processed by an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting and processing a plurality of target sample matrices to be detected based on the data detection method according to any one of claims 1 to 8; According to the prediction result, a target object to be recommended to the target user is determined.

10. A data detection device, the method comprising: An acquisition module, used for acquiring a target sample matrix to be detected; the target sample matrix includes sample features of sample data; the sample data includes voice data, or video data, or text data, and the sample features include voice features, or video features, or text features; A first generating module, used for generating first detection data of the target sample matrix based on position information of the sample feature in the target sample matrix; A second generating module, used for generating second detection data of the target sample matrix based on the eigenvalue of the sample feature; A detection module is used to detect whether the target sample matrix is ​​a sparse sample matrix according to the first detection data and the second detection data.

11. An object recommendation device, comprising: The acquisition module is used to obtain the object operation data and user portrait data of the target user; A prediction module, used for predicting the object operation data and the user portrait data through an object recommendation model to obtain a prediction result; the object recommendation model is obtained by iterative training based on non-sparse target training data; the target training data is obtained by detecting a plurality of target sample matrices to be detected based on the data detection method according to any one of claims 1 to 8; A determination module is used to determine the target object to be recommended to the target user according to the prediction result.

12. An electronic device comprising: processor; as well as, A memory arranged to store computer executable instructions, the executable instructions being configured to be executed by the processor, the executable instructions comprising instructions for executing the steps in the data detection method as described in any one of claims 1 to 8, or the executable instructions comprising instructions for executing the steps in the object recommendation method as described in claim 9.

13. A computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, wherein the executable instructions enable a computer to execute the data detection method according to any one of claims 1 to 8, or the executable instructions include instructions for executing the object recommendation method according to claim 9.

Citation Information

Patent Citations

  • Method for predicting energy consumption of operations of sparse matrix

    CN106547723A

  • Fast sparse neural network

    CN114424252A

  • Feature detection method and device based on sparse matrix vector multiplication and medium

    CN116186526A

  • Data detection method, object recommendation method and device

    CN117951510A

  • Applied estimation of eigenvectors and eigenvalues

    US20050021577A1