<?xml version="1.0" encoding="utf-8"?>
<journal>
<title>Journal of Clinical and Basic Research</title>
<title_fa>Journal of Clinical and Basic Research</title_fa>
<short_title>jcbr</short_title>
<subject>Medical Sciences</subject>
<web_url>http://jcbr.goums.ac.ir</web_url>
<journal_hbi_system_id>1</journal_hbi_system_id>
<journal_hbi_system_user>admin</journal_hbi_system_user>
<journal_id_issn>2538-3736</journal_id_issn>
<journal_id_issn_online>2538-3736</journal_id_issn_online>
<journal_id_pii></journal_id_pii>
<journal_id_doi>10.66224/jcbr</journal_id_doi>
<journal_id_iranmedex></journal_id_iranmedex>
<journal_id_magiran></journal_id_magiran>
<journal_id_sid></journal_id_sid>
<journal_id_nlai></journal_id_nlai>
<journal_id_science></journal_id_science>
<language>en</language>
<pubdate>
	<type>jalali</type>
	<year>1404</year>
	<month>7</month>
	<day>1</day>
</pubdate>
<pubdate>
	<type>gregorian</type>
	<year>2025</year>
	<month>10</month>
	<day>1</day>
</pubdate>
<volume>9</volume>
<number>3</number>
<publish_type>online</publish_type>
<publish_edition>1</publish_edition>
<article_type>fulltext</article_type>
<articleset>
	<article>


	<language>en</language>
	<article_id_doi></article_id_doi>
	<title_fa></title_fa>
	<title>DeepSeek-R1: Reliability for research and education – A comparative study with Claude-3.5-Sonnet and GPT-4o</title>
	<subject_fa>مدیریت آموزشی</subject_fa>
	<subject>Education Management</subject>
	<content_type_fa>پژوهشي</content_type_fa>
	<content_type>Research</content_type>
	<abstract_fa></abstract_fa>
	<abstract>&lt;div style=&quot;text-align: justify;&quot;&gt;&lt;span style=&quot;font-size:12px;&quot;&gt;&lt;span style=&quot;font-family:Times New Roman;&quot;&gt;&lt;b&gt;Background:&lt;/b&gt; Large language models (LLMs) like Claude-3.5-Sonnet and GPT-4o are widely used in research and education but are limited by high costs and proprietary restrictions. DeepSeek-R1, an open-source LLM developed by DeepSeek-AI, leverages a Mixture-of-Experts (MoE) architecture and multi-stage training to offer a cost-effective alternative. This study evaluates DeepSeek-R1&amp;rsquo;s reliability for academic and clinical applications compared to Claude-3.5-Sonnet and GPT-4o, focusing on performance, cost efficiency, and limitations such as censorship and data privacy.&lt;br&gt;
&lt;b&gt;Methods:&lt;/b&gt; A mixed-methods approach was employed, including benchmark evaluations across MATH-500 (mathematics), HumanEval (programming), MMLU (general knowledge), and MedQA (medical reasoning). A prospective user study with 112 Iranian medical researchers assessed diagnostic accuracy on 50 standardized medical cases across specialties (internal medicine, pediatrics, psychiatry). Performance was measured as mean accuracy &amp;plusmn; SD, with paired t-tests (p&lt;0.05) and ANOVA for comparisons. Confidence scores were analyzed using calibration curves (Pearson r). Cost, latency, and limitations (e.g., censorship, data storage) were evaluated using model documentation and reports.&lt;br&gt;
&lt;b&gt;Results:&lt;/b&gt; DeepSeek-R1 achieved 97.3% &amp;plusmn; 1.2 on MATH-500 and 96.3% &amp;plusmn; 1.5 on HumanEval, outperforming Claude-3.5-Sonnet (95.1% &amp;plusmn; 1.4, 94.2% &amp;plusmn; 1.7) and GPT-4o (96.0% &amp;plusmn; 1.3, 95.5% &amp;plusmn; 1.6). MMLU and MedQA accuracies were comparable (90.8% &amp;plusmn; 2.0 and 85.0% &amp;plusmn; 3.2, respectively). In the user study, DeepSeek-R1&amp;rsquo;s diagnostic accuracy (79.2% &amp;plusmn; 4.0) matched Claude-3.5-Sonnet (78.5% &amp;plusmn; 4.2, p=0.42) and GPT-4o (77.8% &amp;plusmn; 4.1, p=0.51), with strong performance in internal medicine (83% &amp;plusmn; 4.5) and pediatrics (81% &amp;plusmn; 5.0). DeepSeek-R1 offered 96% cost savings ($0.14 vs. $4.5/M-tok) and faster latency (42 tokens/s). Limitations include a 4k-token output cap, real-time censorship, and data storage in China.&lt;br&gt;
&lt;b&gt;Conclusion&lt;/b&gt;: DeepSeek-R1 is a reliable, cost-effective alternative to proprietary LLMs, excelling in technical and medical reasoning tasks. Its open-source nature enhances accessibility, but censorship and privacy concerns necessitate careful adoption. Comparative analyses guide its use in academic and clinical settings, emphasizing the need for ethical oversight.&lt;/span&gt;&lt;/span&gt;&lt;/div&gt;</abstract>
	<keyword_fa></keyword_fa>
	<keyword>Artificial Intelligence, Large Language Models, Chatbot, DeepSeek</keyword>
	<start_page>1</start_page>
	<end_page>5</end_page>
	<web_url>http://jcbr.goums.ac.ir/browse.php?a_code=A-10-127-9&amp;slc_lang=en&amp;sid=1</web_url>


<author_list>
	<author>
	<first_name>Reza </first_name>
	<middle_name></middle_name>
	<last_name>Rakhshi </last_name>
	<suffix></suffix>
	<first_name_fa></first_name_fa>
	<middle_name_fa></middle_name_fa>
	<last_name_fa></last_name_fa>
	<suffix_fa></suffix_fa>
	<email>reza.rakhshi21@gmail.com</email>
	<code>10031947532846006576</code>
	<orcid>10031947532846006576</orcid>
	<coreauthor>No</coreauthor>
	<affiliation>Student Research Committee, Golestan University of Medical Sciences, Gorgan, Iran</affiliation>
	<affiliation_fa></affiliation_fa>
	 </author>


	<author>
	<first_name>Teymoor </first_name>
	<middle_name></middle_name>
	<last_name>Khosravi </last_name>
	<suffix></suffix>
	<first_name_fa></first_name_fa>
	<middle_name_fa></middle_name_fa>
	<last_name_fa></last_name_fa>
	<suffix_fa></suffix_fa>
	<email>tkhosravi1375@gmail.com</email>
	<code>10031947532846006577</code>
	<orcid>10031947532846006577</orcid>
	<coreauthor>No</coreauthor>
	<affiliation>Student Research Committee, Golestan University of Medical Sciences, Gorgan, Iran</affiliation>
	<affiliation_fa></affiliation_fa>
	 </author>


	<author>
	<first_name>Arian </first_name>
	<middle_name></middle_name>
	<last_name>Rahimzadeh </last_name>
	<suffix></suffix>
	<first_name_fa></first_name_fa>
	<middle_name_fa></middle_name_fa>
	<last_name_fa></last_name_fa>
	<suffix_fa></suffix_fa>
	<email>Arianrahimzadeh1380@gmail.com</email>
	<code>10031947532846006578</code>
	<orcid>0009-0001-3732-4012</orcid>
	<coreauthor>No</coreauthor>
	<affiliation>Student Research Committee, Golestan University of Medical Sciences, Gorgan, Iran</affiliation>
	<affiliation_fa></affiliation_fa>
	 </author>


	<author>
	<first_name>Mohadeseh </first_name>
	<middle_name></middle_name>
	<last_name>Mohsenipour </last_name>
	<suffix></suffix>
	<first_name_fa></first_name_fa>
	<middle_name_fa></middle_name_fa>
	<last_name_fa></last_name_fa>
	<suffix_fa></suffix_fa>
	<email>mhds.mohsenipour@gmail.com</email>
	<code>10031947532846006579</code>
	<orcid>0009-0004-5079-3288</orcid>
	<coreauthor>No</coreauthor>
	<affiliation>Student Research Committee, Golestan University of Medical Sciences, Gorgan, Iran</affiliation>
	<affiliation_fa></affiliation_fa>
	 </author>


	<author>
	<first_name>Morteza </first_name>
	<middle_name></middle_name>
	<last_name>Oladnabi </last_name>
	<suffix></suffix>
	<first_name_fa></first_name_fa>
	<middle_name_fa></middle_name_fa>
	<last_name_fa></last_name_fa>
	<suffix_fa></suffix_fa>
	<email>oladnabidozin@yahoo.com</email>
	<code>10031947532846006580</code>
	<orcid>10031947532846006580</orcid>
	<coreauthor>Yes
</coreauthor>
	<affiliation>Gorgan Congenital Malformations Research Center, Jorjani Clinical sciences Research Institute, Golestan University of Medical Sciences, Gorgan, Iran; Department of Medical Genetics, School of Advanced Technologies in Medicine, Golestan University of Medical Sciences, Gorgan, Iran</affiliation>
	<affiliation_fa></affiliation_fa>
	 </author>


</author_list>


	</article>
</articleset>
</journal>
