@@ -54,23 +54,23 @@ After the human needs have been successfully assigned for all essays, we evaluat
>> Though the results may again seem utterly dissatisfying and not quite what we had hoped for, it is still somewhat accaptable due to the more difficult and tricky nature of our texts and the novel approach we had to take. The problems that arise from the lack of gold data for certain classes (i.e. "physiological needs") will have to be mitigated either way. A more balanced distribution of texts could possibly solve this problem, but the content of the corpus at use could make this approach quite difficult, as it was the case.
<palign="center">
<imgsrc="pictures/Matrix_test_1.png"/>
<imgsrc="pictures/F1_test.png"/>
</p>
<palign="center">
<imgsrc="pictures/Matrix_test_2.png"/>
</p>
> The system scored a macro-averaged F1 score of roughly 19% for Maslow categories and only 3.5% for Reiss categories. Considering the previously discussed precision and recall measures, these results seem fitting, as the F1 score combines precision and recall and thus serves as a harmonic mean between both measures.
>> The F1 score too may not seem satisfactory, but considering all the circumstantances and factors of the study, F1 score too is still somewhat acceptable, especially for Maslow categories.
> The system scored a macro-averaged F1 score of roughly 19% for Maslow categories and only 3.5% for Reiss categories. Considering the previously discussed precision and recall measures, these results seem fitting, as the F1 score combines precision and recall and thus serves as a harmonic mean between both measures.
>> The F1 score too may not seem satisfactory, but considering all the circumstantances and factors of the study, F1 score too is still somewhat acceptable, especially for Maslow categories.